Text prediction method and device, product and equipment

By generating multiple text unit tokens at each time step in the language model and setting text unit tokens with the same location encoding information in adjacent time steps, the problem of inefficient language model generation is solved, efficient and accurate text generation is achieved, and user experience and model generalization capabilities are improved.

CN120277208APending Publication Date: 2025-07-08TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510352580.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When generating large-scale text data, the method of generating single words one by one is inefficient. How to improve the efficiency of generating the reply content corresponding to the prompt words is a hot topic.

Method used

By gradually performing text prediction processing at each time step, multiple text unit tokens are generated, and a certain number of text unit tokens are set up between adjacent time steps with the same location encoding information, text prediction is used to improve generation efficiency and accuracy.

Benefits of technology

It significantly improves the efficiency and accuracy of text generation, reduces delay, improves user experience, reduces semantic drift, and enhances the generalization ability of language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277208A_ABST
    Figure CN120277208A_ABST
Patent Text Reader

Abstract

The invention discloses a text prediction method and device, a product and equipment. The method comprises the following steps: acquiring prompt information; performing text prediction processing according to at least one time step based on the prompt information to obtain a target prediction text; the text prediction processing corresponding to each time step is used for generating N text units Token associated with each time step, each text unit Token has position coding information, K groups of text units Token with the same position coding information exist between the N text units Token associated with the two adjacent time steps, N is an integer greater than 1, and K is a non-negative integer less than N; generating a reply text corresponding to the prompt information according to the target prediction text; the reply text comprises a text unit Token generated through text prediction processing corresponding to each time step. By adopting the method, the effect of generating the reply text corresponding to the prompt information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of artificial intelligence, and in particular, to a text prediction method, apparatus, product, and device. Background Art

[0002] With the continuous development of AI (Artificial Intelligence), language models have made remarkable progress in the field of natural language processing. By learning context relationships from large-scale text corpora, language models demonstrate powerful language generation and understanding capabilities. After a user inputs a prompt word, the language model can predict and generate corresponding words one by one based on the input prompt word, and the finally generated words can form the reply content corresponding to the prompt word. However, when the data scale of the text to be processed (such as the prompt word) is very large, the efficiency of the language model using this method of generating a single word one by one to generate the reply content of the prompt word is very low. Therefore, how to improve the efficiency of generating the reply content corresponding to the prompt word is a hot issue. Summary of the Invention

[0003] This application provides a text prediction method, apparatus, product, and device, which can improve the efficiency of generating the reply text corresponding to the prompt information.

[0004] On the one hand, this application provides a text prediction method, which includes:

[0005] Obtain prompt information for text prediction;

[0006] Based on the prompt information, perform text prediction processing step by step according to at least one time step to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units (Tokens) associated with each time step respectively, and each text unit (Token) has its own position encoding information. There are K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps. N is an integer greater than 1, and K is a non-negative integer and K is less than N;

[0007] Generate a reply text corresponding to the prompt information according to the target prediction text; the reply text includes the text units (Tokens) generated by the text prediction processing corresponding to each time step.

[0008] On the one hand, this application provides a text prediction apparatus, which includes:

[0009] An acquisition module, configured to obtain prompt information for text prediction;

[0010] A prediction module, configured to perform text prediction processing step by step based on the prompt information according to at least one time step to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units (Tokens) associated with each time step respectively, and each text unit (Token) has its own position encoding information. There are K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps. N is an integer greater than 1, and K is a non-negative integer, and K is less than N;

[0011] A generation module, configured to generate a response text corresponding to the prompt information according to the target prediction text; the response text includes the text units (Tokens) generated by the text prediction processing corresponding to each time step.

[0012] In one implementation, if there is one time step, the N text units (Tokens) associated with one time step are predicted based on a prompt feature sequence, and the prompt feature sequence is generated by performing feature transformation on the prompt information; and,

[0013] If there are multiple time steps, there is a sequential order between the multiple time steps. The N text units (Tokens) associated with any time step are predicted based on the representation features of the text units (Tokens) associated with the time steps before any time step and the prompt feature sequence; wherein, the representation feature of any text unit (Token) is the feature used to decode to obtain any text unit (Token).

[0014] In one implementation, any one of the at least one time step is the t-th time step, and t is a positive integer; the manner in which the prediction module performs text prediction processing step by step based on the prompt information according to at least one time step to obtain a target prediction text includes:

[0015] Obtain a first prediction feature sequence; the first prediction feature sequence is composed of the representation features of the text units (Tokens) in the first prediction text sequence, and the first prediction text sequence is composed of the text units (Tokens) that have been predicted and generated based on the prompt feature sequence in the time steps before the t-th time step;

[0016] According to the first prediction feature sequence and the prompt feature sequence, construct a first reference feature sequence corresponding to the t-th time step;

[0017] Perform text prediction processing using the first reference feature sequence to generate N first text units (Tokens) associated with the t-th time step;

[0018] Obtain a second prediction text sequence based on the N first text units (Tokens), and obtain the target prediction text through the second prediction text sequence.

[0019] In one implementation, the way for the prediction module to obtain the target prediction text from the second prediction text sequence includes:

[0020] If the text unit Token of the end type is not included in the second prediction text sequence, construct a second prediction feature sequence according to the representation features of the text unit Tokens in the second prediction text sequence;

[0021] Construct a second reference feature sequence corresponding to the (t + 1)-th time step according to the second prediction feature sequence and the prompt feature sequence;

[0022] Perform text prediction processing using the second reference feature sequence to generate N second text unit Tokens associated with the (t + 1)-th time step;

[0023] Obtain a third prediction text sequence based on the N second text unit Tokens, and obtain the target prediction text through the third prediction text sequence.

[0024] In one implementation, the prediction module is further configured to:

[0025] If the text unit Token of the end type is included in the second prediction text sequence, use the second prediction text sequence as the target prediction text;

[0026] Wherein, if there is at least one text unit Token after the text unit Token of the end type in the target prediction text, the at least one text unit Token is a zero-valued text unit Token supplemented after the text unit Token of the end type; generating a response text corresponding to the prompt information according to the target prediction text includes:

[0027] Remove the text unit Token of the end type and the zero-valued text unit Tokens in the target prediction text to obtain the response text.

[0028] In one implementation, the position encoding information of any text unit Token is used to indicate the text position of any text unit Token in the response text when the response text uses any text unit Token; the N text positions indicated by the N position encoding information of the N first text unit Tokens are consecutive in sequence, and the N text positions indicated by the N position encoding information of the N second text unit Tokens are consecutive in sequence;

[0029] Among them, there are K groups of text unit Tokens with the same position encoding information among the N first text unit Tokens and the N second text unit Tokens; and, the text position indicated by the position encoding information of the i-th text unit Token among the N first text unit Tokens is before the text position indicated by the position encoding information of the i-th text unit Token among the N second text unit Tokens, where i is a positive integer and i is less than or equal to N.

[0030] In one implementation, the way for the prediction module to obtain the second predicted text sequence based on the N first text unit Tokens includes:

[0031] Obtain t×N text unit Tokens associated with the t-th time step and the time steps before the t-th time step, where the t×N text unit Tokens include N first text unit Tokens;

[0032] Select M reference text unit Tokens corresponding to M position encoding information from the t×N text unit Tokens; the M position encoding information includes the last position encoding information and each position encoding information whose indicated text position is before the text position indicated by the last position encoding information, and the last position encoding information is the position encoding information with the most end text position among the N position encoding information of the N first text unit Tokens, and M is a positive integer;

[0033] Construct the second predicted text sequence using the M reference text unit Tokens corresponding to the M position encoding information.

[0034] In one implementation, any one of the M position encoding information is the target position encoding information, and each text unit Token in the t×N text unit Tokens has its own prediction confidence; the way for the prediction module to select M reference text unit Tokens corresponding to the M position encoding information from the t×N text unit Tokens includes:

[0035] Obtain one or more text unit Tokens with the target position encoding information from the t×N text unit Tokens;

[0036] Take the text unit Token with the highest prediction confidence among the one or more text unit Tokens as the reference text unit Token corresponding to the target position encoding information; or,

[0037] Take the text unit Token with the most occurrences among the one or more text unit Tokens as the reference text unit Token corresponding to the target position encoding information.

[0038] In one implementation, the prediction module uses the first reference feature sequence to perform text prediction processing to generate N first text units (Tokens) associated with the t-th time step, including:

[0039] Obtain the trained language model; the trained language model includes a shared network and N prediction networks;

[0040] Call the shared network to perform feature learning on the first reference feature sequence to generate the shared sequence features of the first reference feature sequence;

[0041] Call each prediction network to perform text prediction processing using the shared sequence features to generate N first text units (Tokens) associated with the t-th time step;

[0042] Wherein, one prediction network is used to generate one first text unit (Token) associated with the t-th time step.

[0043] In one implementation, the target prediction text is generated by calling the trained language model to perform text prediction processing; the above text prediction device further includes a training module, and this training module is used for:

[0044] Obtain the language model and the sample prompt information;

[0045] Call the language model to perform text prediction processing based on the sample prompt information according to the sample time steps to generate N sample text units (Tokens) respectively associated with one or more sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text units (Tokens) respectively associated with each sample time step;

[0046] Based on the N sample text units (Tokens) respectively associated with one or more sample time steps, generate the text prediction loss of the language model for the predicted sample text units (Tokens);

[0047] Use the text prediction loss to correct the model parameters of the language model to obtain the trained language model.

[0048] In one implementation, the N sample text units (Tokens) respectively associated with one or more sample time steps are used to generate the sample reply text corresponding to the sample prompt information. Each sample text unit (Token) associated with one or more sample time steps has its own position encoding information. The position encoding information of any sample text unit (Token) is used to indicate the text position of any sample text unit (Token) in the sample reply text when the sample reply text uses any sample text unit (Token);

[0049] Among them, the sample prompt information has label text, and the label text is a reference for the sample response text. The label text contains L label text unit Tokens, and each of the L label text unit Tokens has position encoding information. The position encoding information of any label text unit Token is used to indicate the text position of any label text unit Token in the label text, and L is a positive integer;

[0050] For any text unit Token associated with any sample time step, the label text unit Token to which the position encoding information of any text unit Token belongs among the L label text unit Tokens is the same, and any text unit Token has a prediction confidence at any sample time step.

[0051] In one implementation, the training module generates the text prediction loss of the language model for the predicted text unit Token based on the N text unit Tokens associated with one or more sample time steps, including:

[0052] Based on the prediction confidences of the N text unit Tokens associated with each sample time step, generate an intermediate prediction loss of the language model at each sample time step;

[0053] Sum up the intermediate prediction losses of the language model at each sample time step to generate a text prediction loss.

[0054] In one implementation, any one of the one or more sample time steps is the s-th sample time step, and s is a positive integer; the training module generates the intermediate prediction loss of the language model at each sample time step based on the prediction confidences of the N text unit Tokens associated with each sample time step, including:

[0055] Based on the prediction confidences of the N text unit Tokens associated with the s-th sample time step, generate a unit prediction loss corresponding to each text unit Token associated with the s-th sample time step;

[0056] Sum up the N unit prediction losses corresponding to the N text unit Tokens associated with the s-th sample time step to generate an intermediate prediction loss of the language model at the s-th sample time step.

[0057] In one implementation, the prompt information is sent by the client; the above-mentioned generation module is further configured to:

[0058] Return the response text to the client, so that the client outputs the response text in the client interface.

[0059] On the one hand, the present application provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the method in one aspect of the present application.

[0060] On the one hand, the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the method in the above-mentioned one aspect.

[0061] On the one hand, the present application provides a computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to execute the method provided in various alternative manners such as the above-mentioned one aspect.

[0062] The present application can obtain hint information for text prediction; and can gradually perform text prediction processing step by step according to at least one time step based on the hint information to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units (Tokens) associated with each time step respectively, each text unit (Token) has its own position encoding information, and there are K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps, N is an integer greater than 1, K is a non-negative integer, and K is less than N; and, a response text corresponding to the hint information can also be generated according to the target prediction text; the response text contains the text units (Tokens) generated by the text prediction processing corresponding to each time step. It can be seen that in the process of gradually performing text prediction processing step by step according to time steps in the method proposed by the present application, multiple text units (Tokens) can be associated and generated at each time step, and there can be K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps, and K can take any non-negative integer less than N. The larger the value of K, the more text units (Tokens) with the same position encoding information can be in the text units (Tokens) associated with two adjacent time steps, the more association features there will be between the front and back time steps during text prediction, and the higher the accuracy of text prediction will be; while the smaller the value of K, the fewer text units (Tokens) with the same position encoding information can be in the text units (Tokens) associated with two adjacent time steps, and the more text units (Tokens) with new position encoding information can be generated at each time step, and the higher the efficiency of text prediction will be. It can be seen that by using the method provided by the present application, by generating multiple text units (Tokens) at each time step and setting a certain number (such as K groups) of text units (Tokens) with the same position encoding information between adjacent time steps, the efficiency and accuracy of text prediction can be improved. Description of the Drawings

[0063] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a schematic structural diagram of the network architecture of a text generation network provided by an embodiment of the present application;

[0065] Figure 2It is a schematic diagram of a scenario for generating a response text corresponding to a prompt message provided by an embodiment of the present application;

[0066] Figure 3 It is a schematic flowchart of a text prediction method provided by an embodiment of the present application;

[0067] Figure 4 It is a schematic diagram of the principle for generating N first text units provided by an embodiment of the present application;

[0068] Figure 5 It is a schematic diagram of the principle for constructing a first reference feature sequence provided by an embodiment of the present application;

[0069] Figure 6 It is a schematic diagram of a scenario for performing text prediction processing based on a prompt message provided by an embodiment of the present application;

[0070] Figure 7 It is a schematic flowchart of a process for training a language model provided by an embodiment of the present application;

[0071] Figure 8 It is a schematic diagram of the principle for training a language model provided by an embodiment of the present application;

[0072] Figure 9 It is a schematic diagram of the structure of a text prediction device provided by an embodiment of the present application;

[0073] Figure 10 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0074] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0075] All data collected in the present application (such as relevant data such as prompt messages and sample prompt messages) are collected with the consent and authorization of the owner of the data (such as users, institutions, or enterprises), and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards in the relevant regions.

[0076] Here, relevant technical concepts involved in the present application are described:

[0077] Token: The smallest unit for text prediction, which can be referred to as a text unit in this application. A text unit is a Token. A Token can be a Chinese character, a word, a character, etc.

[0078] Large Language Model: (abbreviated as LLM), which is a core deep learning technology in the field of AI (Artificial Intelligence). It is a deep learning model trained with a large amount of text data and can generate and understand natural language. In this application, the Large Language Model can be referred to as the language model.

[0079] Time step: A time step represents a single prediction action of the large language model on a Token. That is, a time step can be defined as an independent prediction step when the large language model generates text. This prediction step is used to perform a single prediction and output of a Token.

[0080] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of the network architecture of a text generation network provided by an embodiment of this application. As Figure 1 shown, the network architecture may include a server 200 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices, and the number of terminal devices will not be limited here. As Figure 1 shown, multiple terminal devices may specifically include terminal device 1, terminal device 2, terminal device 3, …, terminal device n; as Figure 1 shown, terminal device 1, terminal device 2, terminal device 3, …, terminal device n can all be network-connected to the server 200, so that each terminal device can perform data interaction with the server 200 through the network connection.

[0081] As Figure 1 shown, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal device can be: smart phones, tablets, laptop computers, desktop computers, smart TVs, in-vehicle terminals, smart homes and other intelligent terminals. Here, taking the communication between terminal device 1 and the server 200 as an example, the specific description of the embodiment of this application will be carried out.

[0082] The terminal device 1 may include a client (such as a Q&A client), and the server 200 may be the background server of the client. It supports that the user can input a prompt message in the client. Here, the prompt message is a prompt text. For example, the prompt text can be a question. After receiving the prompt text, the client can send a reply request to the server 200, and the reply request may carry the prompt text. After receiving the reply request, the server 200 can extract the prompt text from the reply request and generate a reply text corresponding to the prompt text. The reply text is the text used to reply to the prompt text. For example, the reply text can be the answer to the question raised by the prompt text. The server 200 can return the reply text to the client, so that the client can display the reply text on the client interface for the user to view. Among them, for the process of the server 200 generating the reply text corresponding to the prompt text, please refer to the descriptions of the following embodiments.

[0083] Please refer to Figure 2 , Figure 2 is a schematic diagram of a scenario for generating a reply text corresponding to a prompt message provided by an embodiment of the present application. The above Figure 1 After the server 200 in the above obtains the prompt text (i.e., the prompt message), it can perform feature transformation on the prompt text and convert the prompt text into a feature sequence, which can be called a prompt feature sequence. For example, the server 200 can perform word segmentation processing on the prompt text to obtain a prompt text sequence, which is a sequence composed of each word obtained by word segmentation of the prompt text. The server 200 can perform feature learning (such as feature embedding processing) on each word segmentation in the prompt text sequence to obtain a feature vector for each word segmentation. A feature vector can be an embedding (embedded feature). Thus, the feature vectors of each word segmentation can form a prompt feature sequence, which contains the embeddings of each word segmentation in the prompt text. An embedding can be an element in the prompt feature sequence.

[0084] The server 200 can perform text prediction processing step by step according to the time steps through this prompt feature sequence, so as to generate the target prediction text. The process of performing text prediction processing according to the time steps may include text prediction processing corresponding to one or more (here are multiple) time steps. The text prediction processing corresponding to one time step is used to generate multiple text units (which can be called Tokens) associated with this time step. As shown here, through the text prediction processing corresponding to multiple time steps (including time step 1 to time step 6), multiple text units associated with time step 1, multiple text units associated with time step 2, multiple text units associated with time step 3, multiple text units associated with time step 4, multiple text units associated with time step 5, and multiple text units associated with time step 6 can be generated. Among them, the target prediction text can be obtained from the multiple text units associated with each time step. For the specific process of obtaining the target prediction text, reference can also be made to the relevant description in the following Figure 3 corresponding embodiment. Each text unit can have its own position encoding information, and this position encoding information can be used to indicate which word in the text to be finally generated is predicted by the corresponding text unit (that is, which text position in the text to be finally generated is the word predicted). Between the multiple text units (such as N text units, N is a positive integer) respectively associated with adjacent time steps, there can be K groups of text units with the same position encoding information, K is a non-negative integer, and K is less than N.

[0085] Thus, the server 200 can obtain the response text corresponding to the prompt information from the above target prediction text. For example, the useless text unit Tokens in the target prediction text can be removed to obtain the response text. Through the above process, the server 200 can perform text prediction processing according to the time steps through the prompt text, and finally obtain the response text corresponding to the prompt text.

[0086] By using the method provided in this application, at each time step of performing text prediction processing through the prompt information, multiple text units Tokens respectively associated with each time step can be generated. Thus, through the multiple text units Tokens generated and associated with each time step, the response text corresponding to the prompt information can be quickly obtained. Therefore, the efficiency of performing text prediction processing is improved, and the latency of text response to the user is also reduced. Moreover, between the multiple text units Tokens respectively associated with adjacent time steps, there can also be K groups of text units Tokens with the same position encoding information. K is a value adjustable within the range of [0, N), and different values of K can bring different generation effects for the response text. Therefore, by adjusting the value of K in this application, various service requirements for the latency or accuracy of text prediction can also be met. For the specific description of this part of the effect, reference can also be made to the relevant description in step S103 of the following Figure 3 corresponding embodiment.

[0087] Please refer to Figure 3 , Figure 3 which is a schematic flow chart of a text prediction method provided by an embodiment of the present application. The execution subject in the embodiment of the present application may be a text generation device (hereinafter referred to as a generation device for short), and the generation device may be a computer device or a computer device cluster composed of multiple computer devices. The computer device may be a server or other devices, and the present application does not limit this. As Figure 3 shown, the method may include:

[0088] Step S101, obtaining prompt information for text prediction.

[0089] Specifically, the generation device may obtain the prompt information, and the prompt information may be a prompt word (prompt) for a large language model (such as the trained language model described below in the present application). The prompt information may be sent by the client to the generation device, that is, the generation device may obtain the prompt information sent by the client. The generation device may be the background device of the client (such as a background server), and the client may be any client that supports question and answer. For example, the client may be a question and answer client.

[0090] Among them, the prompt information may be any type of information for prompting the large language model to perform language understanding. For example, the prompt information may be text type information (such as prompt text) or / and image type information (such as prompt image), etc. The prompt information may be unimodal information (such as prompt text or prompt image), or may be multimodal information (such as including both prompt text and prompt image), which may be specifically determined according to the actual application scenario. Exemplarily, the prompt text may be any text input by the user on the client. For example, the prompt text may be a question, a word, or any sentence input by the user on the client; the prompt image may be any image uploaded or imported by the user on the client, such as an image of any person, an image of a plant, etc.

[0091] Step S102, performing text prediction processing step by step based on the prompt information according to at least one time step to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units Token associated with each time step respectively, and each text unit Token has its own position encoding information. There are K groups of text unit Tokens with the same position encoding information between the N text unit Tokens associated with two adjacent time steps. N is an integer greater than 1, and K is a non-negative integer and K is less than N.

[0092] Specifically, the generation device can perform text prediction processing progressively according to the above prompt information in at least one time step to generate a target prediction text. The target prediction text can be the prediction text ultimately used to generate the response text corresponding to the prompt information. The target prediction text can be the last prediction text sequence before the end of the text prediction processing. For specific details, please refer to the following description.

[0093] Among them, the process of the generation device performing text prediction processing according to time steps can include text prediction processing corresponding to one or more time steps (i.e., at least one time step). The text prediction processing corresponding to each time step can be respectively used to generate N text units (Tokens) associated with each time step. N is an integer greater than 1. That is, when performing text prediction processing at any time step, multiple text units (Tokens) associated with that time step can be generated (such as N text units (Tokens)). The specific value of N can be flexibly set according to actual business requirements. For example, N can be set to 3, 4, 5, etc.

[0094] Among them, each text unit (Token) can have its own position encoding information. One text unit (Token) associated with a time step can have one position encoding information. The position encoding information of any text unit (Token) can be used to indicate the text position (such as the nth word in the response text) where the text unit (Token) is located in the response text when the text unit (Token) is used in the subsequent response text.

[0095] There can be K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps (which can be any two adjacent time steps). K is a non - negative integer and K is less than N. The value of K can be set according to actual business requirements. K can be any non - negative integer less than N, that is, K can be any integer in [0, N). In other words, there may or may not be text units (Tokens) with overlapping text positions between the N text units (Tokens) predicted by adjacent time steps.

[0096] The generation device can obtain a trained language model, and the trained language model can be a large - language model (such as an LLM model). The process of training the trained language model can be found in the specific description in the corresponding embodiment below. Figure 7 The generation device can call the trained language model to perform text prediction processing step by step according to the prompt information in at least one time step to generate a target prediction text. Among them, the trained language model can be used to predict and generate one text unit (Token) at a time step.

[0097] If there exists a time step, the N text units Tokens associated with this time step can be predicted through a prompt feature sequence, and this prompt feature sequence is generated by performing feature transformation on the prompt information.

[0098] If there are multiple time steps, there is a sequential order among these multiple time steps. The text prediction process corresponding to any one time step can be carried out based on the text prediction process corresponding to the previous time step. That is, the N text units Tokens associated with any one time step can be predicted through the representation features of the text units Tokens associated with the time step before this time step and the above-mentioned prompt feature sequence. The text prediction processes corresponding to each time step can be carried out sequentially in order. A text unit can be a Token generated by prediction (the smallest unit for text prediction), such as a character or a phrase that can be predicted.

[0099] Among them, the representation feature of any one text unit Token is the feature used to decode this text unit Token, and this feature can be the feature generated and output by the prediction network in the following language model.

[0100] In one implementation, the process of performing feature transformation on the prompt information to generate a prompt feature sequence may include: the feature encoding (i.e., feature transformation) of the prompt information can be realized through an information encoder adapted to the information type of the prompt information to obtain the prompt feature sequence of the prompt information.

[0101] For example, if the prompt information is a prompt text, each token in this prompt text can be subjected to feature encoding processing through a text encoder to generate the feature vector (embedding) of each token. Thus, the prompt feature sequence of this prompt text can be formed through the feature vectors of each token.

[0102] Another example, if the prompt information is a prompt image, and there is one or more such prompt images, each prompt image can be subjected to feature encoding processing through an image encoder to generate the feature vector (embedding) of each prompt image. Thus, the corresponding prompt feature sequence can also be formed through the feature vectors of each prompt image.

[0103] Still another example, if the prompt information includes both the above-mentioned prompt text and prompt images at the same time, the feature sequence obtained by performing feature transformation on the prompt text and the feature sequence obtained by performing feature transformation on the prompt images can be concatenated to obtain the prompt feature sequence. This prompt feature sequence can include the feature vectors of each token in the prompt text and the feature vectors of each prompt image.

[0104] The above information encoder (such as a text encoder or an image encoder) can be a pre-trained model capable of feature encoding (i.e., feature transformation) of information of a suitable information type, or the information encoder can also be part of a trained language model and is obtained by training together with the trained language model.

[0105] If the information encoder is a model outside the trained language model, the above prompt feature sequence can be used as the input to the trained language model; if the information encoder is part of the trained language model, the above prompt information can be used as the input to the trained language model.

[0106] In one implementation, any one of the above at least one time step can be called the t-th time step, where t is a positive integer, that is, the t-th time step can be any time step in the process of text prediction processing. The following takes the process of text prediction processing corresponding to the t-th time step as an example for specific description. It can be understood that the principle of text prediction processing corresponding to each time step is the same, but the text prediction processing corresponding to each time step can be gradually carried out in sequence according to the order of the time steps.

[0107] The generation device can obtain a first predicted feature sequence, which can be composed of the representation features of text units (Tokens) in the first predicted text sequence, that is, the first predicted feature sequence can include the representation features of each text unit (Token) in the first predicted text sequence. The first predicted text sequence can be composed of text units (Tokens) that have been predicted and generated through the prompt feature sequence at time steps before the t-th time step. If t is equal to 1, there are no other time steps before the t-th time step, and the first predicted text sequence can be empty, that is, there may be no first predicted text sequence, and thus no first predicted feature sequence, because there are no text units (Tokens) that have been predicted and generated through the prompt feature sequence before the t-th time step. That is, when t is equal to 1, there may be no first predicted feature sequence, but N text units (Tokens) associated with the first time step predicted through the prompt feature sequence. If t is greater than 1, the principle of obtaining the first predicted text sequence is the same as the principle of obtaining the following second predicted text sequence, and the specific principle of obtaining the second predicted text sequence can be referred to the following description.

[0108] The generating device can construct a reference feature sequence corresponding to the t-th time step according to the first predicted feature sequence and the above-mentioned prompt feature sequence, and the reference feature sequence corresponding to the t-th time step can be called the first reference feature sequence. For example, the first reference feature sequence can be a feature sequence obtained by concatenating the first predicted feature sequence and the prompt feature sequence, that is, the first predicted feature sequence and the prompt feature sequence can be concatenated (the concatenating order can be a preset order) to obtain the first reference feature sequence. The first reference feature sequence can be the input at the t-th time step, that is, at the t-th time step, text prediction can be performed through the first reference feature sequence, as described below.

[0109] The generating device can call the trained language model to perform text prediction processing according to the first reference feature sequence to generate N text units Token associated with the t-th time step, and the N text units Token associated with the t-th time step can be called N first text units Token.

[0110] The trained language model in this application can include a shared network and N prediction networks. The shared network can also be called the backbone network of the trained language model, and the N prediction networks can share the backbone network. The generating device can input the above-mentioned first reference feature sequence into the shared network to call the shared network to perform feature learning (which can be embedding processing) on the first reference feature sequence to generate the shared sequence feature of the first reference feature sequence. The shared sequence feature is the feature learned by the shared network from the first reference feature sequence (belonging to implicit feature representation). The backbone network can be a complex Transformer network, and each prediction network can be a lightweight Transformer layer. The model parameters of the backbone network can be much more than the model parameters of each prediction network. Therefore, from the overall perspective of the language model, the introduction of multiple prediction networks in this application hardly changes the number of model parameters of the language model, and the model parameters of each prediction network are a very small part of the model parameters of the entire language model.

[0111] The generation device can call each prediction network in the trained language model to perform text prediction processing using the shared sequence feature, so as to generate N first text units Token associated with the t-th time step. One prediction network can be used to perform text prediction processing using the shared sequence feature to generate one first text unit Token associated with the t-th time step. That is, the number of prediction networks can be equal to N. In actual business requirements, if it is necessary to generate a certain number of text units Token associated with one time step, then the same number of prediction networks can be set in the language model. The N prediction networks can be used to quickly generate the N text units Token associated with the time step in parallel at one time step.

[0112] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the principle of generating N first text units provided by an embodiment of the present application. As Figure 4 shown, the trained language model can include a shared network and N prediction networks. Here, N can be equal to 4, that is, there are 4 prediction networks, including prediction network 1 to prediction network 4. The generation device can input the first reference feature sequence into the shared network to call the shared network to perform feature learning on the first reference feature sequence and generate the shared sequence feature (embedded feature) of the first reference feature sequence. The shared sequence feature can be input into each prediction network, and each prediction network can perform text prediction processing using the shared sequence feature to generate the corresponding predicted text feature, which is the feature output by the prediction network performing text prediction processing using the shared sequence feature. Here, prediction network 1 can generate predicted text feature 1, prediction network 2 can generate predicted text feature 2, prediction network 3 can generate predicted text feature 3, and prediction network 4 can generate predicted text feature 4.

[0113] The text unit Token1 can be mapped through the predicted text feature 1, the text unit Token2 can be mapped through the predicted text feature 2, the text unit Token3 can be mapped through the predicted text feature 3, and the text unit Token4 can be mapped through the predicted text feature 4. The N first text units Token associated with the t-th time step can include the text unit Token1, text unit Token2, text unit Token3, and text unit Token4. Among them, for the principle of mapping and decoding the predicted text feature obtained by the prediction network to obtain the corresponding text unit Token, please refer to the principle described in formula (1) below.

[0114] Among them, the above prediction text feature 1 can be the representation feature of text unit Token1, prediction text feature 2 can be the representation feature of text unit Token2, prediction text feature 3 can be the representation feature of text unit Token3, and prediction text feature 4 can be the representation feature of text unit Token4. That is, the representation feature of any text unit is the feature generated by the prediction network for decoding to obtain this text unit, and this feature can be an embedding feature.

[0115] The generating device can obtain a second predicted text sequence through the above N first text unit Tokens, and can obtain a target predicted text through this second predicted text sequence. The nature of this second predicted text sequence is the same as the concept of the nature of the above first predicted text.

[0116] If the second predicted text sequence contains the predicted text unit Token of the end type (i.e., the end symbol), it indicates that the process of text prediction processing ends, and there is no need to perform text prediction processing corresponding to the next time step. The second predicted text sequence can be directly used as the target predicted text.

[0117] If the second predicted text sequence does not contain the text unit of this end type, it indicates that the process of text prediction processing has not ended, and text prediction processing corresponding to the next time step needs to be performed. Therefore, the generating device can continue to construct a reference feature sequence corresponding to the (t + 1)-th time step according to the representation features of the text unit Tokens in the second predicted text sequence and the above hint feature sequence. For example, the generating device can construct a second predicted feature sequence according to the representation features of the text unit Tokens in the second predicted text sequence. That is, this second predicted feature sequence is composed of the representation features of each text unit Token in the second predicted text sequence. The reference feature sequence corresponding to the (t + 1)-th time step can be called the second reference feature sequence. Similarly to the above first reference feature sequence, this second reference feature sequence can also be obtained by concatenating the second predicted feature sequence and the hint feature sequence. The generating device can call the trained language model to perform text prediction processing through this second reference feature sequence to generate N text unit Tokens associated with the (t + 1)-th time step. The N text unit Tokens associated with the (t + 1)-th time step can be called N second text unit Tokens. Among them, the principle of performing text prediction processing through the second reference feature sequence to generate N second text unit Tokens is the same as the principle of performing text prediction processing through the first reference feature sequence to generate N first text unit Tokens.

[0118] In addition, the generation device can obtain a third predicted text sequence through the N second text unit Tokens, and continue to generate a target predicted text through the third predicted text sequence. Among them, the principle of generating the target predicted text through the third predicted text sequence is the same as the principle of generating the target predicted text through the second predicted text sequence above. The main thing is to determine whether the text unit Token of the end type is included in the currently generated predicted text sequence, that is, to determine whether the process of text prediction processing is over. If it is not over, continue with the text prediction processing corresponding to the next time step. If it is over, the currently generated predicted text sequence can be used as the finally generated target predicted text.

[0119] Both the above-mentioned N first text unit Tokens and N second text unit Tokens can have their own position encoding information, that is, each text unit Token can have its own position encoding information, and a text unit Token can have one position encoding information. The position encoding information of any text unit Token can be used to indicate the text position of the any text unit in the subsequent reply text when the any text unit is adopted. In other words, the position encoding information of a text unit Token is used to indicate which word the text unit Token is predicted for in the reply text. Simply put, the position encoding information of a text unit Token is used to indicate the text position of the text unit Token, and this text position can be understood as the position number of the text unit Token, such as the position number being which word the text unit Token is predicted for in the reply text.

[0120] The N text positions indicated by the N position encoding information of the N first text unit Tokens are sequentially continuous, that is, the N text positions are adjacent to each other in sequence. In addition, the N text positions indicated by the N position encoding information of the N second text unit Tokens are also sequentially continuous, that is, the N text positions are also adjacent to each other in sequence. That is to say, the N text positions indicated by the N position encoding information of the N text unit Tokens generated at any time step are all sequentially continuous and adjacent, that is, the N text unit Tokens generated at any time step can be N words with continuous positions generated at the any time step.

[0121] In one implementation, both the above-mentioned N first text unit Tokens and N second text unit Tokens can also have their own position encoding information, and there can be K groups of text unit Tokens with the same position encoding information among the N first text unit Tokens and N second text unit Tokens.

[0122] Moreover, the text position indicated by the position encoding information of the $i$-th text unit Token among the $N$ first text unit Tokens (sorted in the order of text positions) is before the text position indicated by the position encoding information of the $i$-th text unit Token among the $N$ second text unit Tokens, where $i$ is a positive integer and $i \leq N$.

[0123] That is to say, the overall text position of the $N$ text unit Tokens associated with the previous time step is before the overall text position of the $N$ text unit Tokens associated with the next time step. For example, the overall text position of the $N$ first text unit Tokens associated with the $t$-th time step is before the overall text position of the $N$ second text unit Tokens associated with the $(t + 1)$-th time step. That is, the next time step predicts and generates at least one (one or more) new text unit Tokens at subsequent text positions compared to the previous time step.

[0124] For example, the above $N$ first text unit Tokens may include the first text unit $w1$ at the predicted first text position, the first text unit $w2$ at the second text position, the first text unit $w3$ at the third text position, and the first text unit $w4$ at the fourth text position; and the above $N$ second text unit Tokens may include the second text unit $w5$ at the predicted third text position, the second text unit $w6$ at the fourth text position, the second text unit $w7$ at the fifth text position, and the second text unit $w8$ at the sixth text position.

[0125] Therefore, there can be 2 groups of text units with the same position encoding information among the $N$ first text unit Tokens and the $N$ second text unit Tokens, that is, $K = 2$. One group of text unit Tokens with the same position encoding information may include the first text unit $w3$ and the second text unit $w5$, which have the same position encoding information indicating the third text position; and another group of text unit Tokens with the same position encoding information may include the first text unit $w4$ and the second text unit $w6$, which have the same position encoding information indicating the fourth text position.

[0126] Another example, the above $N$ first text unit Tokens may include the first text unit $w1$ at the predicted first text position, the first text unit $w2$ at the second text position, and the first text unit $w3$ at the third text position; and the above $N$ second text unit Tokens may include the second text unit $w4$ at the predicted fourth text position, the second text unit $w5$ at the fifth text position, and the second text unit $w6$ at the sixth text position.

[0127] Therefore, there can be 0 groups of text unit tokens with the same position encoding information among the N first text unit tokens and the N second text unit tokens, that is, K is equal to 0. At this time, there are no text unit tokens with the same position encoding information among the N first text unit tokens and the N second text unit tokens.

[0128] From the above, the process of obtaining the second predicted text sequence from the above-mentioned N first text unit tokens can include: the generating device can obtain the t-th time step and the t×N text unit tokens associated with the time steps before the t-th time step (if any). That is, there are currently t time steps in total, and one time step has N associated text unit tokens, so t time steps are associated with t×N text unit tokens. The t×N text unit tokens can include the N first text unit tokens associated with the t-th time step.

[0129] The generating device can select, from the t×N text unit tokens, the reference text unit tokens corresponding to M position encoding information respectively. The M position encoding information can include the last position encoding information and each position encoding information indicating that the text position is before the text position indicated by the last position encoding information. The last position encoding information can be the position encoding information indicating the text position at the last position among the N position encoding information of the N first text unit tokens. M is a positive integer.

[0130] For example, in the order of the text positions, the above-mentioned N first text unit tokens can successively include the first text unit token at the predicted second text position, the first text unit token at the third text position, the first text unit token at the fourth text position, and the first text unit token at the fifth text position. Then, the position encoding information indicating the fifth text position can be the last position encoding information, and the M position encoding information can include this last position encoding information, the position encoding information indicating the fourth text position, the position encoding information indicating the third text position, the position encoding information indicating the second text position, and the position encoding information indicating the first text position. In this case, the N text unit tokens associated with the time step before the t-th time step can include the text unit tokens at the predicted first text position, the second text position, the third text position, and the fourth text position.

[0131] Any one of the above M position encoding information can be called target position encoding information. Since the principle of obtaining the reference text unit Token corresponding to each position encoding information is the same, the following will take obtaining the reference text unit corresponding to the target position encoding information as an example for specific description. One position encoding information can correspond to one reference text unit Token.

[0132] The generating device can obtain one or more text unit Tokens with target position encoding information from the above t×N text unit Tokens. Among them, if there are multiple time steps and there are text unit Tokens with overlapping text positions among the N text unit Tokens associated with different time steps, there can be multiple text unit Tokens with target position encoding information; if there is only one time step, or there are no text unit Tokens with overlapping text positions among the N text unit Tokens associated with different time steps, there can be only one text unit Token with target position encoding information.

[0133] Among them, each of the above t×N text unit Tokens can have its own prediction confidence. One text unit Token can have one prediction confidence. The prediction confidence of a text unit Token is the confidence obtained when predicting the text unit Token at the corresponding time step. The prediction confidence of a text unit Token can also be the prediction probability of the text unit Token. The higher the prediction confidence, the greater the probability that the trained language model believes that the text unit Token is at the corresponding text position; conversely, the lower the prediction confidence, the smaller the probability that the trained language model believes that the text unit Token is at the corresponding text position.

[0134] Therefore, the generating device takes the text unit Token with the highest prediction confidence among the one or more text unit Tokens with target position encoding information as the reference text unit Token corresponding to the target position encoding information. Or, the generating device can take the text unit Token with the highest occurrence frequency (i.e., the number of occurrences, which is the number) among the one or more text unit Tokens as the reference text unit Token corresponding to the target position encoding information. How to specifically select the reference text unit Token corresponding to the target position encoding information can be flexibly set according to the actual application scenario.

[0135] The generating device may use the M reference text unit Tokens corresponding to the M position encoding information obtained above to construct a second predicted text sequence, and the second predicted text sequence may include the M reference text unit Tokens. The arrangement order of the M reference text unit Tokens in the second predicted text sequence may be the sequential order of the text positions indicated by the position encoding information corresponding to the M reference text unit Tokens.

[0136] As can be seen from the above, the second predicted text sequence is the optimal text sequence verified from the text unit Tokens generated at the previous time steps (such as verified by prediction confidence). This text sequence may be composed of the optimal text unit Tokens at each text position that has been predicted currently. For example, the text unit Token with a higher prediction confidence can be considered as the optimal text unit Token, or the text unit Token with a higher occurrence frequency at the same text position can be considered as the optimal text unit Token at this text position.

[0137] Among them, since when generating text unit Tokens at each time step, N text unit Tokens with consecutive text positions are generated, therefore, in this application, the order of text positions when the above N prediction networks generate text unit Tokens can be set. For example, the N prediction networks may include Prediction Network 1, Prediction Network 2, Prediction Network 3, and Prediction Network 4. Then, for the 4 consecutive text positions where text unit Tokens need to be generated currently, it can be set that Prediction Network 1 is used to generate the text unit Token at the 1st text position among the 4 text positions, Prediction Network 2 is used to generate the text unit Token at the 2nd text position among the 4 text positions, Prediction Network 3 is used to generate the text unit Token at the 3rd text position among the 4 text positions, and Prediction Network 4 is used to generate the text unit Token at the 4th text position among the 4 text positions.

[0138] Specifically, there may also be a set order among the N prediction networks. The i-th prediction network among the N prediction networks may be used to generate the text unit Token at the i-th text position among the N consecutive text positions where text unit Tokens need to be generated currently. i is a positive integer less than or equal to N.

[0139] Among them, the i-th prediction network may be denoted as f i , the above backbone network may be denoted as Z zg , the above first reference feature sequence may be denoted as x ck, through this backbone network and the i-th prediction network, at the t-th time step, among the N consecutive text positions where text unit Tokens need to be generated, at the text position ranked i-th, the prediction confidence for each candidate text unit Token in the candidate text unit library (i.e., the vocabulary library, which can also be simply referred to as the word library and can contain a large number of candidate text unit Tokens, i.e., a large number of words) can be generated. The prediction confidence for each candidate text unit Token can together form a probability distribution. Finally, the candidate text unit Token with the highest prediction confidence in the candidate text unit library can be taken as the text unit Token generated at the t-th time step and located at the text position ranked i-th. Therefore, at the text position ranked i-th, a probability distribution P for the candidate text units in the candidate text unit library is generated. fb The principle can be shown as the following formula:

[0140] P fb = softmax(f u (f i (Z zg (x ck )))) (1)

[0141] Among them, P fb can include the prediction confidence for each candidate text unit Token in the candidate text unit library at the text position ranked i-th. softmax is a normalization function. f u can be the word embedding decoding matrix (which can be simply referred to as the decoding matrix or can also be called the decoding network) in the trained language model. This word embedding decoding matrix can be used to map the feature representation of the hidden layer (such as the features output by the prediction network) to the dimension of the vocabulary (i.e., the vocabulary library). This decoding matrix belongs to the model parameters of the language model. Each prediction network can share this decoding matrix for feature decoding processing to obtain the predicted text unit Token.

[0142] As can be understood from the above, when predicting a text unit Token at any time step, the N position encoding information of N consecutive text positions that need to be text-predicted currently can be known in advance. Therefore, based on the fact that the N position encoding information can be obtained in advance, the text unit Tokens corresponding to the N position encoding information can be predicted and generated respectively, so as to generate the N text unit Tokens associated with any time step. The text unit Token corresponding to a position encoding information can have the position encoding information, and the position encoding information is used to indicate the text position corresponding to the text unit Token in the text that finally needs to be generated, that is, to indicate that if the text unit Token is used in the text that finally needs to be generated (such as the following reply text), which word the text unit Token is in the text that finally needs to be generated.

[0143] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the principle of constructing a first reference feature sequence provided by an embodiment of the present application. As Figure 5 shown, the above-mentioned t-th time step can be the 4th time step, and there can also be the 1st time step to the 3rd time step before the 4th time step. Assume that N is equal to 3 and K is equal to 2. The N text unit Tokens generated by the text prediction process corresponding to the 1st time step include text unit 1, text unit 2, and text unit 3; the N text unit Tokens generated by the text prediction process corresponding to the 2nd time step include text unit 11, text unit 22, and text unit 33; the N text unit Tokens generated by the text prediction process corresponding to the 3rd time step include text unit 111, text unit 222, and text unit 333. That is, there can be two sets of text unit Tokens with the same position encoding information between the N text unit Tokens associated with any two adjacent time steps.

[0144] Here, for example, text unit 2 and text unit 11 can have the same position encoding information; text unit 3, text unit 22, and text unit 111 can have the same position encoding information; and text unit 33 and text unit 222 can have the same position encoding information.

[0145] Among them, the position encoding information of text unit 1 is position encoding information 1, and text unit 1 is predicted from the 1st word; the position encoding information of text units 2 and 11 is position encoding information 2, and text units 2 and 11 are predicted from the 2nd word; the position encoding information of text units 3, 22, and 111 is position encoding information 3, and text units 3, 22, and 111 are predicted from the 3rd word; the position encoding information of text units 33 and 222 is position encoding information 4, and text units 33 and 222 are predicted from the 4th word; and the position encoding information of text unit 333 is position encoding information 5, and text unit 333 is predicted from the 5th word. It shows that there are a total of words predicted from 5 text positions before the 4th time step, and these 5 text positions are the text positions of the first 5 words.

[0146] Therefore, if the prediction confidence of text unit 2 in text units 2 and 11 is the highest, the prediction confidence of text unit 22 in text units 3, 22, and 111 is the highest, and the prediction confidence of text unit 33 in text units 33 and 222 is the highest, then the reference text unit corresponding to position encoding information 1 can be text unit 1, the reference text unit corresponding to position encoding information 2 can be text unit 2, the reference text unit corresponding to position encoding information 3 can be text unit 22, the reference text unit corresponding to position encoding information 4 can be text unit 33, and the reference text unit corresponding to position encoding information 5 can be text unit 333.

[0147] The above first predicted text sequence can be composed of the reference text unit Token corresponding to the position encoding information 1 (i.e., text unit 1), the reference text unit Token corresponding to the position encoding information 2 (i.e., text unit 2), the reference text unit Token corresponding to the position encoding information 3 (i.e., text unit 22), the reference text unit Token corresponding to the position encoding information 4 (i.e., text unit 33), and the reference text unit Token corresponding to the position encoding information 5 (i.e., text unit 333). And by concatenating the first predicted feature sequence formed by the representation features of each text unit Token in the first predicted text sequence with the above prompt feature sequence, the input at the 4th time step, that is, the above first reference feature sequence corresponding to the 4th time step, can be obtained.

[0148] At the 4th time step, text prediction processing can be performed through the first reference feature sequence to continue generating N text units (Tokens) associated with the 4th time step. And so on, the N text units (Tokens) associated with the 4th time step can include the text unit (Token) predicted for the 4th word, the text unit (Token) predicted for the 5th word, and the text unit (Token) predicted for the 6th word.

[0149] Step S103, generate a response text corresponding to the prompt information according to the target prediction text; the response text contains the text units (Tokens) generated by the text prediction processing corresponding to each time step.

[0150] Specifically, if there are at least one other text unit (Token) after the text unit (Token) of the above end type in the target prediction text, the at least one text unit (Token) can be a zero-valued text unit (Token) (i.e., 0 element) supplemented by the trained language model after the text unit (Token) of the end type during text prediction processing. The text unit (Token) of the end type is usually an indicator (i.e., end symbol) indicating the end of prediction, and the text unit (Token) of the end type usually may not be reflected in the final response text.

[0151] Therefore, the generating device can perform a removal process (i.e., remove) on the text unit (Token) of the end type in the target prediction text and the zero-valued text unit (Token) supplemented after the text unit (Token) of the end type, and then the response text can be obtained. The response text can contain the text units (Tokens) generated by the text prediction processing corresponding to each of the above time steps. In the response text, there can be a text unit (Token) at a text position for prediction.

[0152] The response text can be the text obtained by integrating in sequence (in the order of the sequence of the text positions corresponding to each text unit) the text units (Tokens) in the target prediction text except for the end symbol and the zero-valued text unit (Token). The integration together can mean putting together (such as connecting together). The response text is the text generated by the trained language model through text prediction processing for the prompt information to reply to the prompt information. For example, if the prompt information is a question raised by the user, the response text can be an answer to the question; another example is that if the prompt information is a vocabulary input by the user, the response text can be a text explaining and describing the meaning and related events of the vocabulary; and so on.

[0153] The generating device can return the above reply text to the client, enabling the client to output the reply text in the client interface for the user to view and use, such as supporting the user to forward, comment on, provide feedback on, share the reply text, and so on.

[0154] In one implementation effect, the larger the value of K, the more groups of text unit Tokens with overlapping text positions among the N text unit Tokens generated by adjacent time steps, enabling the language model to learn more features related to the association between the front and back text positions. Therefore, the accuracy of the language model for text prediction can be higher. If the value of K is smaller, the fewer groups of text unit Tokens with overlapping text positions among the N text unit Tokens generated by adjacent time steps, enabling the language model to generate text unit Tokens at more text positions with fewer time steps, reducing the total number of steps (i.e., the number of total time steps) required for the language model to perform text prediction. Therefore, the efficiency of the language model for text prediction can be higher. Therefore, in an actual business scenario, the value of K can be set by comprehensively considering the accuracy and efficiency of text prediction according to the actual business requirements. The role of the value of K can also be reflected in the following training process of the language model. The larger the value of K, the more groups of sample text unit Tokens with the same label among the N sample text unit Tokens generated by adjacent sample time steps.

[0155] Among them, K can be any integer between [0, N). The larger K is, the more association features between the front and back text positions the language model learns, the fewer text unit Tokens the language model generates for a new text position in one time step, and the slower the language model's thinking speed. Therefore, the strategy of taking a larger value of K can be called a slow thinking strategy. During the slow thinking strategy process, the accuracy of the text unit Tokens predicted by the language model for each text position can be higher because the text unit Tokens at the same text position can be gradually and continuously optimized at different time steps. The smaller K is, the fewer association features between the front and back text positions the language model learns, the more text unit Tokens the language model generates for a new text position in one time step, and the faster the language model's thinking speed. Therefore, the strategy of taking a smaller value of K can be called a fast thinking strategy.

[0156] In one embodiment, when the requirement for latency is higher than the requirement for accuracy, the smaller the value of K is taken, the higher the efficiency of text prediction will be. For example, the value of K can be taken within the first value range, such as the first value range can be [0, N / 2), that is, the value of K can be taken as a value less than N / 2. At this time, it can be understood that the above-mentioned fast thinking strategy is adopted; when the requirement for accuracy is higher than the requirement for latency, the larger the value of K is taken, the higher the accuracy of text prediction will be. For example, the value of K can be taken within the second value range, and the second value range can be (N / 2, N), that is, the value of K can be taken as a value greater than N / 2. At this time, it can be understood that the above-mentioned slow thinking strategy is adopted. The values within the first value range can be generally smaller than the values within the second value range.

[0157] Among them, the specific value of K (that is, whether to adopt the slow thinking strategy or the fast thinking strategy) can be flexibly adjusted and changed according to the business requirements of the actual business scenario. For example, the specific value of K can be set according to the relevant business attributes of text prediction, and the business attributes can be business attributes related to the latency or accuracy rate of text prediction. For example, when the network quality between the generating device and the client is poor, in order to reply to the user faster and avoid more latency, the fast thinking strategy can be adopted for text prediction; and in the Q&A scenarios in some more professional technical fields, when the requirement for the accuracy of the replied content is relatively high, the slow thinking strategy can be adopted for text prediction. And so on.

[0158] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a scenario for text prediction processing through prompt information provided by an embodiment of the present application. Such as Figure 6As shown, the prompt message can be prompt text (which can be tokenized to include the words "knowledge", "is", "power", and the added end symbol "Esep"). The shared network can include a text encoder. The prompt text can be input into the shared network (which can be denoted as Shared Trunk) to call the shared network to perform feature transformation on the prompt text, obtaining the prompt feature sequence of the prompt text. Also, the shared network can be called to perform feature learning on the prompt feature sequence (which can be at the first time step), and the features generated by the shared network during feature learning (such as the above-mentioned shared sequence features) can be input into each prediction network for text prediction processing. Among them, the prediction network can also be called an independent prediction head, which can be denoted as Independent Output Head. Here, there are a total of 4 independent prediction heads, that is, N takes the value of 4. These 4 independent prediction heads can include Independent Output Head1, Independent Output Head2, Independent Output Head3, and Independent Output Head4.

[0159] Exemplarily, if the above-mentioned slow thinking strategy is adopted, here, through the text prediction network of each prediction network, text units Token at the text positions indicated by the position encoding information T1 (which can be the position of the first word), text units Token at the text positions indicated by the position encoding information T2 (which can be the position of the second word), text units Token at the text positions indicated by the position encoding information T3 (which can be the position of the third word), and text units Token at the text positions indicated by the position encoding information T4 (which can be the position of the fourth word) can be generated at the first time step; and text units Token at the text positions indicated by the position encoding information T2 (which can be the position of the second word), text units Token at the text positions indicated by the position encoding information T3 (which can be the position of the third word), text units Token at the text positions indicated by the position encoding information T4 (which can be the position of the fourth word), and text units Token at the text positions indicated by the position encoding information T5 (which can be the position of the fifth word) can be generated at the second time step; and, text units Token at the text positions indicated by the position encoding information T3 (which can be the position of the third word), text units Token at the text positions indicated by the position encoding information T4 (which can be the position of the fourth word), text units Token at the text positions indicated by the position encoding information T5 (which can be the position of the fifth word), and text units Token at the text positions indicated by the position encoding information T6 (which can be the position of the sixth word) can be generated at the third time step. And so on, more text units Token associated with different time steps can be generated. Finally, when the text unit Token of the end type is predicted, the process of text prediction processing can be ended.

[0160] Among them, in the above-mentioned slow thinking strategy, K is taken as the maximum value 3 within its value range [0, N), and there can be 3 groups of text units Token with the same position encoding information between 4 text units Token associated with two adjacent time steps respectively. Through the respective text units Token generated at each time step above, the extremely accurate response text corresponding to the prompt text can finally be obtained according to the principle described above.

[0161] Moreover, if the above-mentioned fast thinking strategy is adopted, the text prediction network of each prediction network here can generate text tokens at the text positions indicated by the position encoding information T1 (which can be the position of the first word), text tokens at the text positions indicated by the position encoding information T2 (which can be the position of the second word), text tokens at the text positions indicated by the position encoding information T3 (which can be the position of the third word), and text tokens at the text positions indicated by the position encoding information T4 (which can be the position of the fourth word) at the first time step; and can generate text tokens at the text positions indicated by the position encoding information T5 (which can be the position of the fifth word), text tokens at the text positions indicated by the position encoding information T6 (which can be the position of the sixth word), text tokens at the text positions indicated by the position encoding information T7 (which can be the position of the seventh word), and text tokens at the text positions indicated by the position encoding information T8 (which can be the position of the eighth word) at the second time step; and can generate text tokens at the text positions indicated by the position encoding information T9 (which can be the position of the ninth word), text tokens at the text positions indicated by the position encoding information T10 (which can be the position of the tenth word), text tokens at the text positions indicated by the position encoding information T11 (which can be the position of the eleventh word), and text tokens at the text positions indicated by the position encoding information T12 (which can be the position of the twelfth word) at the third time step. And so on, more text tokens associated with different time steps can be generated. Finally, when the text token of the end type is predicted, the process of text prediction processing can be ended.

[0162] Among them, the above-mentioned fast thinking strategy sets K to the minimum value 0 within its value range [0, N). There are 0 groups of text tokens with the same position encoding information between the 4 text tokens associated with two adjacent time steps, that is, there is no group of text tokens with the same position encoding information between the 4 text tokens associated with two adjacent time steps. Through the text tokens generated at each time step as described above, the response text corresponding to the prompt text can be obtained efficiently and quickly according to the principle described above.

[0163] This application proposes an objective of multi-token prediction, which can predict multiple future tokens simultaneously at each time step. Through a shared backbone network and N independent output head structures (i.e., prediction networks, also known as prediction heads), the language model can learn and capture more global semantic information of the context through multiple tokens associated with each previous time step at each time step, reducing the over-reliance of the language model on historical generation results (such as tokens at previous text positions), thereby reducing the cumulative error in the text generation process of the language model. Therefore, the accuracy of text prediction to generate text unit tokens can be greatly improved. Moreover, since this application can predict multiple future tokens synchronously and in parallel at each time step, the speed of text prediction to generate text unit tokens can also be greatly improved, thus enhancing the efficiency of text prediction.

[0164] Based on this, in natural language generation tasks, by adopting the multi-token prediction of this application, due to learning and capturing more global semantic information of the context, the phenomenon of semantic drift in text generation is also reduced, the coherence and consistency of the text content in the generated text are improved, and the generalization ability of the language model for text prediction is enhanced.

[0165] For users, the method provided by this application can generate a response text corresponding to the user's prompt information with extremely low latency, achieving a fast and efficient response to the user and enhancing the user experience. The latency of this application in the text generation task of online services, compared with using traditional language models (such as a language model that generates one token per time step), has been significantly optimized from 2.6 seconds to 1.37 seconds, and the latency has been basically reduced by half, demonstrating excellent text generation performance and greatly enhancing the efficiency of text generation.

[0166] The present application can obtain hint information for text prediction; and can gradually perform text prediction processing step by step according to at least one time step based on the hint information to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units (Tokens) respectively associated with each time step, each text unit (Token) has its own position encoding information, and there are K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) respectively associated with two adjacent time steps, N is an integer greater than 1, K is a non-negative integer, and K is less than N; and, it can also generate a reply text corresponding to the hint information according to the target prediction text; the reply text contains the text units (Tokens) generated by the text prediction processing corresponding to each time step. Thus, it can be seen that in the process of gradually performing text prediction processing step by step according to the time step in the method proposed by the present application, multiple text units (Tokens) can be associated and generated at each time step, and there can be K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) respectively associated with two adjacent time steps, and K can take any non-negative integer less than N. The larger the value of K, the more text units (Tokens) with the same position encoding information can be in the text units (Tokens) respectively associated with two adjacent time steps, and the more association features there will be between the front and back time steps during text prediction, and the higher the accuracy of text prediction will be; while the smaller the value of K, the fewer text units (Tokens) with the same position encoding information can be in the text units (Tokens) respectively associated with two adjacent time steps, and the more text units (Tokens) with new position encoding information can be generated at each time step, and the higher the efficiency of text prediction will be. It can be seen that by adopting the method provided by the present application, by generating multiple text units (Tokens) at each time step and setting a certain number (such as K groups) of text units (Tokens) with the same position encoding information between adjacent time steps, the efficiency and accuracy of text prediction can be improved.

[0167] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a process for training a language model provided by an embodiment of the present application. The above-mentioned target prediction text can be generated by calling the trained language model for text prediction processing. Therefore, the embodiment of the present application specifically describes the process of how to train a language model to obtain the trained language model. As Figure 7 shown, the process may include:

[0168] Step S201, obtain a language model and sample hint information.

[0169] Specifically, the generation device can obtain a language model and sample prompt information. The concept of the sample prompt information can be the same as the concept of the above-mentioned prompt information, and the sample prompt information is a sample for training the language model.

[0170] The sample prompt information can have label text, which can be set for the sample prompt information and is the ideal sample response text to be generated through the sample prompt information. The relationship between the sample response text and the sample prompt information is similar to the relationship between the above-mentioned response text and the prompt information, and the sample response text can be a text for replying to the sample prompt information.

[0171] Step S202: Call the language model to perform text prediction processing based on the sample prompt information according to the sample time steps, and generate N sample text units (Tokens) associated with each of one or more sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text units (Tokens) associated with each of the respective sample time steps.

[0172] Specifically, the generation model can call the language model to perform text prediction processing according to the sample prompt information according to the sample time steps to generate N sample text units (Tokens) associated with each of one or more sample time steps. The concept of the sample time step is similar to the concept of the above-mentioned time step, but here it is called the sample time step for the sake of distinguishing the training process.

[0173] Among them, the principle of calling the language model to perform text prediction processing according to the sample prompt information according to the sample time steps to generate N sample text units (Tokens) associated with each of one or more sample time steps is the same as the principle of calling the trained language model to perform text prediction processing according to the prompt information according to the time steps to generate N text units (Tokens) associated with each of one or more time steps. The text prediction processing process also ends when a text unit (Token) of the end type is predicted. The one or more sample time steps are the respective sample time steps before the end of the text prediction processing process. Among them, how many sample time steps are specifically performed before the end of the text prediction processing process can be determined according to the actual result of the language model performing text prediction processing.

[0174] Similarly, the process of performing text prediction processing according to the sample time steps can also include the text prediction processing corresponding to the one or more sample time steps, and the text prediction processing corresponding to each sample time step is used to generate N sample text units (Tokens) associated with each of the respective sample time steps.

[0175] The N sample text unit Tokens associated with each of the one or more sample time steps can be used to generate a text prediction loss of the language model for the predicted text unit Token. This text prediction loss can be used to reflect the prediction deviation of the language model for the predicted sample text unit Token at each sample time step. The greater the text prediction loss, the greater the prediction deviation; the smaller the text prediction loss, the smaller the prediction deviation.

[0176] Similarly, the N sample text unit Tokens associated with each of the one or more sample time steps can be used to generate a sample response text corresponding to the sample prompt information. The principle of generating the sample response text through the N sample text unit Tokens associated with each sample time step is the same as the principle of generating the response text through the N text unit Tokens associated with each time step described above. For example, it can include the relevant principles of obtaining the above-mentioned second predicted text sequence.

[0177] The above-mentioned tag text can be used as a reference for the sample response text. That is, the most ideal situation is when the sample response text is the same as the tag text sequence. In the actual training process, the sample response text may not be generated. The introduction of the sample response text here is to illustrate the position encoding information of each sample text unit Token.

[0178] The concept of the position encoding information of the sample text unit Token can be similar to the concept of the position encoding information of the above-mentioned text unit Token. Each of the N sample text unit Tokens associated with the one or more sample time steps can have its own position encoding information. For example, the position encoding information of any sample text unit Token can also be used to indicate the text position of the any sample text unit Token in the sample response text when the any sample text unit Token is used in the sample response text. In other words, the position encoding information of a sample text unit Token can be used to indicate which word the sample text unit Token is predicted to be. There can also be K groups of the position encoding information of the sample text unit Tokens that are the same (i.e., overlapping) between the N sample text unit Tokens associated with two adjacent sample time steps.

[0179] Among them, the above label text may contain L label text unit Tokens, and each of the L label text unit Tokens can have position encoding information. The position encoding information of any one label text unit Token can be used to indicate the text position of the any one label text unit Token in the label text. For example, the position encoding information of the any one label text unit Token can be used to indicate which word the any one label text unit Token is in the label text. One label text unit Token is the most ideal text unit Token to be predicted at the text position indicated by the position encoding information of the label text unit Token. For example, if the position encoding information of the label text unit Token indicates that the label text unit Token is the 3rd word in the label text, it means that when performing text prediction processing, it is the most ideal situation that the 3rd word is predicted and generated as the label text unit Token. L is a positive integer, and the specific value of L can be determined according to the actual application scenario. L can be the total number of text unit Tokens (such as the text unit Tokens that make up the sample response text) that finally need to be predicted, generated, and adopted.

[0180] Any one sample text unit Token associated with any one sample time step can be the same as the label text unit Token to which the position encoding information of the any one sample text unit Token belongs among the L label text unit Tokens. The any one sample text unit Token can have a prediction confidence at the any one sample time step. In other words, during the training process, when generating sample text unit Tokens at each sample time step, it can be based on the position encoding information of the currently to-be-generated (i.e., needed to be generated) sample text unit Token to select the label text unit Token with the same position encoding information from the L label text unit Tokens as the currently predicted and generated sample text unit Token.

[0181] Among them, when generating a sample text unit Token corresponding to any one position encoding information at a sample time step, prediction confidences for each candidate text unit Token in the candidate text unit library (i.e., the vocabulary library, which can be abbreviated as the word library) will be obtained. Therefore, at this time, the candidate text unit Token in the candidate text unit library that is the same as the label text unit Token to which the position encoding information belongs can be used as the predicted and generated sample text unit Token. The prediction confidence of the sample text unit Token is not necessarily the largest among the prediction confidences of each candidate text unit Token in the candidate text unit library.

[0182] For example, a total of 3 sample time steps can exist. The N sample text units Token associated with the first sample time step can include: the sample text unit y1 at the first text position (i.e., the first word predicted overall at the first sample time step), the sample text unit y2 at the second text position (i.e., the second word predicted overall at the first sample time step), and the sample text unit y3 at the third text position (i.e., the third word predicted overall at the first sample time step).

[0183] The N sample text units Token associated with the second sample time step can include: the sample text unit y11 at the third text position (i.e., the third word predicted overall at the second sample time step), the sample text unit y22 at the fourth text position (i.e., the fourth word predicted overall at the second sample time step), and the sample text unit y33 at the fifth text position (i.e., the fifth word predicted overall at the second sample time step).

[0184] Similarly, the N sample text units Token associated with the third sample time step can include: the sample text unit y111 at the fifth text position (i.e., the fifth word predicted overall at the third sample time step), the sample text unit y222 at the sixth text position (i.e., the sixth word predicted overall at the third sample time step), and the sample text unit y333 at the seventh text position (i.e., the seventh word predicted overall at the third sample time step).

[0185] In the above example, there can be 1 group of sample text unit Tokens with the same position encoding information (i.e., the same text position) between any two adjacent sample time steps associated with the N sample text unit Tokens, that is, K equals 1. For example, these two adjacent sample time steps can include the first sample time step and the second sample time step, and the second sample time step and the third sample time step.

[0186] If the above sample text unit y222 is an end - type text unit Token, then the label text can include a total of 5 label text unit Tokens from the first text position to the fifth text position. These 5 label text unit Tokens can include: the label text unit b1 at the first text position, the label text unit b2 at the second text position, the label text unit b3 at the third text position, the label text unit b4 at the fourth text position, and the label text unit b5 at the fifth text position.

[0187] Therefore, the above sample text unit y1 can be the label text unit b1, the sample text unit y2 can be the label text unit b2, the sample text unit y3 can be the label text unit b3, the sample text unit y11 can be the label text unit b3, the sample text unit y22 can be the label text unit b4, the sample text unit y33 can be the label text unit b5, and the sample text unit y111 can also be the label text unit b5. Since the sample text unit y222 is an end-type text unit Token, the sample text unit y333 can be a zero-value text unit (i.e., the supplemented 0 element).

[0188] Among them, each sample text unit Token associated with each sample time step has its own prediction confidence at the associated sample time step.

[0189] Step S203: Generate the text prediction loss of the language model for the predicted sample text unit Token based on the N sample text unit Tokens associated with one or more sample time steps.

[0190] Specifically, the generating device can generate the text prediction loss of the language model for the predicted sample text unit Token as a whole through the N sample text unit Tokens associated with each of the above sample time steps, as described below.

[0191] The generating device can generate the intermediate prediction loss of the language model at each sample time step through the prediction confidence of the N sample text unit Tokens associated with each sample time step. The language model can have an intermediate prediction loss at a sample time step, and this intermediate prediction loss reflects the prediction deviation of the N sample text unit Tokens associated with this sample time step.

[0192] Any one of the above one or more sample time steps can be called the s-th sample time step, where s is a positive integer, that is, the s-th sample time step can be any one of the above sample time steps. Since the principle of generating the intermediate prediction loss of the language model at each sample time step through the prediction confidence of the N sample text unit Tokens associated with each sample time step is the same, here, taking the process of generating the intermediate prediction loss of the language model at the s-th sample time step through the prediction confidence of the N sample text unit Tokens associated with the s-th sample time step as an example, a specific description is given.

[0193] The generation device can generate the unit prediction loss corresponding to each sample text unit token associated with the s-th sample time step based on the prediction confidence of each of the N sample text unit tokens associated with the s-th sample time step. Thus, the N unit prediction losses corresponding to the N sample text unit tokens associated with the s-th sample time step can be summed up (i.e., summed) to generate the intermediate prediction loss of the language model at the s-th sample time step.

[0194] In one implementation, based on a conditional independence assumption, the present application can decompose and represent the distribution of the N prediction confidences (i.e., prediction probabilities) of the N sample text unit tokens associated with the s-th sample time step as the product of the prediction confidences of these N sample text unit tokens, as shown in the following formula:

[0195]

[0196] where P θ can be the product of the N prediction confidences of the N sample text unit tokens associated with the s-th sample time step, and P i can represent the prediction confidence of the i-th sample text unit token among the N sample text unit tokens, and i is a positive integer less than or equal to N.

[0197] Therefore, the intermediate prediction loss L s of the language model at the s-th sample time step can be:

[0198]

[0199] In the above formula, log represents taking the logarithm, and log(P i ) represents the unit prediction loss corresponding to the i-th sample text unit token, and L s is the intermediate prediction loss of the language model at the s-th sample time step.

[0200] The generation device can, according to the above principle, obtain the intermediate prediction losses of the language model at each sample time step (all sample time steps experienced at the end of the text prediction process). The generation device can sum up the intermediate prediction losses of the language model at each sample time step (i.e., sum), and then the above text prediction loss of the language model as a whole can be generated.

[0201] Therefore, assuming there are z sample time steps in total, where z is a positive integer, the final text prediction loss L z of the language model as a whole can be:

[0202]

[0203] Among them, j represents the j-th sample time step, where j is a positive integer less than or equal to z, and L z represents the final text prediction loss of the language model. This text prediction loss reflects the overall prediction deviation of the language model for each sample text unit Token associated with each sample time step.

[0204] Step S204: Use the text prediction loss to correct the model parameters of the language model to obtain the trained language model.

[0205] Specifically, the generating device can use the obtained text prediction loss to correct the model parameters of the language model, and finally obtain the above-mentioned trained language model. This trained language model can be applied to the actual text prediction processing scenario for text generation, such as generating the response text corresponding to the above prompt information.

[0206] Among them, the goal of using the text prediction loss to correct the model parameters of the language model can be to make the text prediction loss tend to be minimized (such as tending to 0), so that when the subsequent language model generates each sample text unit Token at each sample time step, the prediction confidence in the corresponding label text unit at the text position can be higher.

[0207] The above sample prompt information can be a large amount. According to the same principle above, the present application can perform multiple rounds of iterative training on the language model through a large amount of sample prompt information. After the training is completed, the above-mentioned trained language model can be obtained. For example, the completion of the training can refer to training the model parameters of the language model to a convergent state, or referring to the number of training rounds of the language model reaching a set threshold number of rounds.

[0208] Please refer to Figure 8 , Figure 8 which is a schematic diagram of the principle of training a language model provided by an embodiment of the present application. As Figure 8 shown, here, the language model can include 1 shared network and 3 prediction networks (including prediction network 1 to prediction network 3). And, there can be a total of 3 sample time steps. Through the language model, N sample text unit Tokens associated with each sample time step can be obtained, including N sample text units associated with the first sample time step, N sample text units associated with the second sample time step, and N sample text units associated with the third sample time step.

[0209] Through the N sample text units associated with the first sample time step, the intermediate prediction loss 1 of the language model at the first sample time step can be generated; through the N sample text units associated with the second sample time step, the intermediate prediction loss 2 of the language model at the second sample time step can be generated; through the N sample text units associated with the third sample time step, the intermediate prediction loss 3 of the language model at the third sample time step can be generated. Thus, by summing up the intermediate prediction loss 1, the intermediate prediction loss 2, and the intermediate prediction loss 3, the overall text prediction loss of the language model can be obtained. The text prediction loss can be backpropagated to the language model to correct the model parameters of the language model, and finally the above-trained language model can be obtained.

[0210] Through multi-Token prediction, this application can not only predict the next Token (such as the Token at the next text position), but also predict multiple future Tokens, significantly improving the language model's ability to model complex context patterns, enhancing the language model's ability to capture long-distance dependencies and global semantic information, enabling the language model to focus on future contexts at longer distances (such as more distant text positions). Multi-Token prediction can effectively capture global dependencies at the sentence, paragraph, or even discourse level, thereby improving the training efficiency of the language model and the utilization rate of the language model for samples (such as sample prompt information), and greatly enhancing the training effect of the language model.

[0211] Through the construction of the multi-Token prediction target, this application has achieved significant improvements to the traditional language model in terms of both the prediction target and the model structure (such as introducing N independent prediction networks), providing a reliable theoretical and technical basis for efficient model training and high-quality text generation.

[0212] In the actual experimental process, the present application also used 4 prediction heads (i.e., 4 prediction networks) to verify the above method proposed by the present application. The speed of text prediction has increased by 1.5 to 3 times compared to traditional language models, and the accuracy of text prediction has been significantly improved. For traditional language models, each time step generates one Token, and the text generation time increases linearly. The text generation complexity is O(T), where T is the length of the generated text sequence. In contrast, the present application generates multiple Tokens at each time step, and the text generation complexity becomes O(T / N), resulting in a significant reduction in text generation time. Moreover, in the experimental process of code generation (i.e., the text to be generated is code), when using the language model of the present application compared to using traditional language models, the accuracy of code generation on the HumanEval test set (a dataset of programming questions) has increased by 7%, and the accuracy of code generation on the MBPP test set (a dataset of programming problems) has also increased by 12%. This demonstrates the effectiveness and practicality of the above method proposed by the present application.

[0213] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a text prediction device provided by an embodiment of the present application. As Figure 9 shown, the text prediction device 90 may include: an acquisition module 901, a prediction module 902, and a generation module 903.

[0214] The acquisition module 901 is configured to acquire prompt information for text prediction;

[0215] The prediction module 902 is configured to gradually perform text prediction processing based on the prompt information according to at least one time step to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text unit Tokens associated with each time step respectively, and each text unit Token has its own position encoding information. There are K groups of text unit Tokens with the same position encoding information between the N text unit Tokens associated with two adjacent time steps. N is an integer greater than 1, and K is a non-negative integer, and K is less than N;

[0216] The generation module 903 is configured to generate a response text corresponding to the prompt information according to the target prediction text; the response text includes the text unit Tokens generated by the text prediction processing corresponding to each time step.

[0217] In one implementation, if there is one time step, the N text unit Tokens associated with one time step are predicted based on a prompt feature sequence, and the prompt feature sequence is generated by performing feature transformation on the prompt information; and,

[0218] If there are multiple time steps, there is a sequential order among the multiple time steps. The N text units Token associated with any time step are predicted based on the representation features and prompt feature sequences of the text units Token associated with the time steps before any time step; among them, the representation feature of any text unit Token is the feature used to decode any text unit Token.

[0219] In one implementation, any one of at least one time step is the t-th time step, where t is a positive integer; the prediction module 902 performs text prediction processing step by step according to at least one time step based on the prompt information, and the manner of obtaining the target prediction text includes:

[0220] Obtain the first prediction feature sequence; the first prediction feature sequence is composed of the representation features of the text units Token in the first prediction text sequence, and the first prediction text sequence is composed of the text units Token that have been predicted and generated based on the prompt feature sequence in the time steps before the t-th time step;

[0221] According to the first prediction feature sequence and the prompt feature sequence, construct the first reference feature sequence corresponding to the t-th time step;

[0222] Use the first reference feature sequence to perform text prediction processing to generate N first text units Token associated with the t-th time step;

[0223] Obtain the second prediction text sequence based on the N first text units Token, and obtain the target prediction text through the second prediction text sequence.

[0224] In one implementation, the manner in which the prediction module 902 obtains the target prediction text through the second prediction text sequence includes:

[0225] If the second prediction text sequence does not contain text units Token of the end type, construct the second prediction feature sequence according to the representation features of the text units Token in the second prediction text sequence;

[0226] According to the second prediction feature sequence and the prompt feature sequence, construct the second reference feature sequence corresponding to the (t + 1)-th time step;

[0227] Use the second reference feature sequence to perform text prediction processing to generate N second text units Token associated with the (t + 1)-th time step;

[0228] Obtain the third prediction text sequence based on the N second text units Token, and obtain the target prediction text through the third prediction text sequence.

[0229] In one implementation, the prediction module 902 is further configured to:

[0230] If the second predicted text sequence contains a text unit Token of the end type, then use the second predicted text sequence as the target predicted text;

[0231] Wherein, if there is at least one text unit Token after the text unit Token of the end type in the target predicted text, then at least one text unit Token is a zero-valued text unit Token supplemented after the text unit Token of the end type; generating a response text corresponding to the prompt information according to the target predicted text, including:

[0232] Remove the text unit Token of the end type and the zero-valued text unit Token in the target predicted text to obtain the response text.

[0233] In one implementation, the position encoding information of any text unit Token is used to indicate the text position of any text unit Token in the response text when the response text adopts any text unit Token; the N text positions indicated by the N position encoding information of the N first text unit Tokens are successively continuous, and the N text positions indicated by the N position encoding information of the N second text unit Tokens are successively continuous;

[0234] Wherein, there are K groups of text unit Tokens with the same position encoding information among the N first text unit Tokens and the N second text unit Tokens; and, the text position indicated by the position encoding information of the i-th text unit Token among the N first text unit Tokens is located before the text position indicated by the position encoding information of the i-th text unit Token among the N second text unit Tokens, where i is a positive integer and i is less than or equal to N.

[0235] In one implementation, the manner in which the prediction module 902 obtains the second predicted text sequence based on the N first text unit Tokens includes:

[0236] Obtain the t×N text unit Tokens associated with the t-th time step and the time steps before the t-th time step, and the t×N text unit Tokens include N first text unit Tokens;

[0237] Select the reference text unit Tokens corresponding to M position encoding information from the t×N text unit Tokens; the M position encoding information includes the last position encoding information and each position encoding information whose indicated text position is before the text position indicated by the last position encoding information, and the last position encoding information is the position encoding information with the text position located at the last among the N position encoding information of the N first text unit Tokens, and M is a positive integer;

[0238] Construct a second predicted text sequence using M reference text unit Tokens corresponding to M position encoding information.

[0239] In one implementation, any one of the M position encoding information is target position encoding information, and each of the t×N text unit Tokens has its own prediction confidence; the prediction module 902 selects the reference text units corresponding to the M position encoding information from the t×N text unit Tokens in the following ways:

[0240] Obtain one or more text unit Tokens with the target position encoding information from the t×N text unit Tokens;

[0241] Take the text unit Token with the highest prediction confidence among the one or more text unit Tokens as the reference text unit Token corresponding to the target position encoding information; or,

[0242] Take the text unit Token with the most occurrences among the one or more text unit Tokens as the reference text unit Token corresponding to the target position encoding information.

[0243] In one implementation, the prediction module 902 uses the first reference feature sequence for text prediction processing to generate N first text unit Tokens associated with the t-th time step in the following ways:

[0244] Obtain the trained language model; the trained language model includes a shared network and N prediction networks;

[0245] Call the shared network to perform feature learning on the first reference feature sequence to generate the shared sequence feature of the first reference feature sequence;

[0246] Call each prediction network to perform text prediction processing using the shared sequence feature to generate N first text unit Tokens associated with the t-th time step;

[0247] Among them, one prediction network is used to generate one first text unit Token associated with the t-th time step.

[0248] In one implementation, the target predicted text is generated by calling the trained language model for text prediction processing; the above text prediction device 90 further includes a training module 904, and the training module 904 is used for:

[0249] Obtain the language model and the sample prompt information;

[0250] Call the language model to perform text prediction processing based on the sample prompt information according to the sample time steps, and generate N sample text unit Tokens associated with each of one or more sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text unit Tokens associated with each of the sample time steps.

[0251] Based on the N sample text unit Tokens associated with each of one or more sample time steps, generate the text prediction loss of the language model for the predicted sample text unit Tokens.

[0252] Use the text prediction loss to correct the model parameters of the language model to obtain the trained language model.

[0253] In one implementation, the N sample text unit Tokens associated with each of one or more sample time steps are used to generate the sample response text corresponding to the sample prompt information. Each sample text unit Token associated with one or more sample time steps has its own position encoding information. The position encoding information of any sample text unit Token is used to indicate the text position of the sample text unit Token in the sample response text when the sample response text uses the sample text unit Token.

[0254] Among them, the sample prompt information has a label text, and the label text is a reference for the sample response text. The label text contains L label text unit Tokens. The L label text unit Tokens all have position encoding information. The position encoding information of any label text unit Token is used to indicate the text position of the label text unit Token in the label text, where L is a positive integer.

[0255] Any sample text unit Token associated with any sample time step is the same as the label text unit Token to which the position encoding information of the sample text unit Token belongs among the L label text unit Tokens. Any sample text unit Token has a prediction confidence level at any sample time step.

[0256] In one implementation, the manner in which the training module 904 generates the text prediction loss of the language model for the predicted sample text unit Tokens based on the N sample text unit Tokens associated with each of one or more sample time steps includes:

[0257] Based on the prediction confidence levels of the N sample text unit Tokens associated with each sample time step, generate the intermediate prediction loss of the language model at each sample time step.

[0258] Perform a summation process on the intermediate prediction losses of the language model at each sample time step to generate the text prediction loss.

[0259] In one embodiment, any one of one or more sample time steps is the s-th sample time step, where s is a positive integer; the training module 904 generates the intermediate prediction loss of the language model at each sample time step based on the prediction confidence of the N sample text units (Tokens) associated with each sample time step, and the method includes:

[0260] Based on the prediction confidence of each of the N sample text units (Tokens) associated with the s-th sample time step, generate a unit prediction loss corresponding to each sample text unit (Token) associated with the s-th sample time step;

[0261] Sum up the N unit prediction losses corresponding to the N sample text units (Tokens) associated with the s-th sample time step to generate the intermediate prediction loss of the language model at the s-th sample time step.

[0262] In one embodiment, the prompt information is sent by the client; the above-mentioned generation module 903 is further configured to:

[0263] Return the reply text to the client, so that the client outputs the reply text in the client interface.

[0264] According to an embodiment of the present application, Figure 3 the steps involved in the text prediction method shown can be Figure 9 executed by each module in the text prediction device 90 shown. For example, Figure 3 the step S101 shown in Figure 9 can be executed by the acquisition module 901 in Figure 3 the step S102 shown in Figure 9 can be executed by the prediction module 902 in Figure 3 the step S103 shown in Figure 9 can be executed by the generation module 903 in

[0265] The present application can obtain hint information for text prediction; and can perform text prediction processing step by step according to at least one time step based on the hint information to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units (Tokens) associated with each time step respectively, each text unit (Token) has its own position encoding information, and there are K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps, N is an integer greater than 1, K is a non-negative integer, and K is less than N; and, a response text corresponding to the hint information can also be generated according to the target prediction text; the response text includes the text units (Tokens) generated by the text prediction processing corresponding to each time step. It can be seen that in the process of performing text prediction processing step by step according to time steps, the method proposed in the present application can generate multiple text units (Tokens) associated with each time step, and there can be K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps, and K can take any non-negative integer less than N. The larger the value of K, the more text units (Tokens) with the same position encoding information can be in the text units (Tokens) associated with two adjacent time steps, and the more association features there will be between the previous and subsequent time steps during text prediction, and the higher the accuracy of text prediction will be; while the smaller the value of K, the fewer text units (Tokens) with the same position encoding information can be in the text units (Tokens) associated with two adjacent time steps, and the more text units (Tokens) with new position encoding information can be generated at each time step, and the higher the efficiency of text prediction will be. It can be seen that by using the method provided in the present application, by generating multiple text units (Tokens) at each time step and setting a certain number (such as K groups) of text units (Tokens) with the same position encoding information between adjacent time steps, the efficiency and accuracy of text prediction can be improved.

[0266] According to an embodiment of the present application, Figure 9 Each module in the text prediction device 90 shown can be separately or entirely combined into one or several units to form, or a certain one (or some) of the units can be further split into multiple smaller sub-units in terms of function, and the same operation can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be realized by multiple units, or the functions of multiple modules can be realized by one unit. In other embodiments of the present application, the text prediction device 90 can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0267] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0268] According to an embodiment of the present application, a computer program capable of executing the steps involved in the corresponding methods shown in the embodiments of the present application can be run on a general-purpose computer device (which may include processing elements such as a central processing unit (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), etc. and storage elements) to construct a text prediction device 90 as shown in Figure 9 shown. The above computer program can be recorded on a computer-readable recording medium, and can be loaded into the above computer device through the computer-readable recording medium and run therein.

[0269] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As shown in Figure 10 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, in some embodiments, the computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and optionally the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory, for example, at least one disk memory. The memory 1005 may optionally be at least one storage device located far from the aforementioned processor 1001. As shown in Figure 10 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0270] In Figure 10In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for users to input; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to achieve:

[0271] Obtain hint information for text prediction;

[0272] Based on the hint information, perform text prediction processing step by step according to at least one time step to obtain the target prediction text; the text prediction processing corresponding to each time step is used to generate N text units Token associated with each time step respectively, and each text unit Token has its own position encoding information. There are K groups of text unit Tokens with the same position encoding information between the N text unit Tokens associated with two adjacent time steps. N is an integer greater than 1, and K is a non-negative integer, and K is less than N;

[0273] Generate a response text corresponding to the hint information according to the target prediction text; the response text contains the text unit Tokens generated by the text prediction processing corresponding to each time step.

[0274] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the above text prediction method in the embodiments of the present application, and can also execute the description of the above text prediction device 90 in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 9 The description of the above text prediction device 90 in the corresponding embodiments described above will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0275] In addition, it should be pointed out here that: the present application also provides a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium. When the processor executes the computer program, it can execute the description of the text prediction method in the embodiments of the present application. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer storage medium involved in the present application, please refer to the description of the method embodiments of the present application.

[0276] As an example, the above computer program can be deployed to be executed on a computer device, or be deployed to be executed on multiple computer devices located at one place, or, be executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain network.

[0277] The above computer-readable storage medium may be an internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0278] The present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, enabling the computer device to execute the descriptions of the above text prediction method in the embodiments of the present application. Therefore, the descriptions will not be repeated here. In addition, the descriptions of the beneficial effects of the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the descriptions of the method embodiments of the present application.

[0279] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products or equipment.

[0280] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0281] The above disclosure is only the preferred embodiment of the present application. Of course, it cannot be used to limit the scope of the rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A text prediction method, characterized in that, The method includes: Obtaining hint information for text prediction; Based on the hint information, performing text prediction processing step by step according to at least one time step to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units (Tokens) associated with each time step respectively, each text unit (Token) has its own position encoding information, and there are K groups of text units (Tokens) with the same position encoding information between the N text units (Tokens) associated with two adjacent time steps, N is an integer greater than 1, K is a non - negative integer, and K is less than N; Generating a reply text corresponding to the hint information according to the target prediction text; the reply text contains the text units (Tokens) generated by the text prediction processing corresponding to each time step.

2. The method according to claim 1, wherein If there is one time step, the N text units (Tokens) associated with the one time step are predicted based on a hint feature sequence, and the hint feature sequence is generated by performing feature transformation on the hint information; and, If there are multiple time steps, there is a sequential order among the multiple time steps, and the N text units (Tokens) associated with any one time step are predicted based on the representation features of the text units (Tokens) associated with the time steps before any one time step and the hint feature sequence; wherein, the representation feature of any one text unit (Token) is the feature used to decode to obtain any one text unit (Token).

3. The method according to claim 2, wherein Any one of at least one time step is the t - th time step, t is a positive integer; the performing text prediction processing step by step according to at least one time step based on the hint information to obtain a target prediction text includes: Obtaining a first prediction feature sequence; the first prediction feature sequence is composed of the representation features of the text units (Tokens) in a first prediction text sequence, and the first prediction text sequence is composed of the text units (Tokens) predicted and generated based on the hint feature sequence in the time steps before the t - th time step; According to the first prediction feature sequence and the hint feature sequence, constructing a first reference feature sequence corresponding to the t - th time step; Performing text prediction processing using the first reference feature sequence to generate N first text units (Tokens) associated with the t - th time step; Obtaining a second prediction text sequence based on the N first text units (Tokens) and obtaining the target prediction text through the second prediction text sequence.

4. The method according to claim 3, characterized in that, The obtaining the target prediction text through the second prediction text sequence includes: If the second prediction text sequence does not contain text units (Tokens) of an end type, constructing a second prediction feature sequence according to the representation features of the text units (Tokens) in the second prediction text sequence; According to the second prediction feature sequence and the hint feature sequence, constructing a second reference feature sequence corresponding to the (t + 1)-th time step; Performing text prediction processing using the second reference feature sequence to generate N second text units (Tokens) associated with the (t + 1)-th time step; Obtain a third predicted text sequence based on the N second text unit Tokens, and obtain the target predicted text through the third predicted text sequence.

5. The method according to claim 4, characterized in that, The method further includes: If the second predicted text sequence contains the text unit Token of the end type, use the second predicted text sequence as the target predicted text; Wherein, if there is at least one text unit Token after the text unit Token of the end type in the target predicted text, the at least one text unit Token is a zero-valued text unit Token supplemented after the text unit Token of the end type; The generating the response text corresponding to the prompt information according to the target predicted text includes: Remove the text unit Token of the end type and the zero-valued text unit Token in the target predicted text to obtain the response text.

6. The method according to claim 4, wherein The position encoding information of any one of the text unit Tokens is used to indicate the text position of any one of the text unit Tokens in the response text when the response text uses any one of the text unit Tokens; the N text positions indicated by the N position encoding information of the N first text unit Tokens are consecutive in sequence, and the N text positions indicated by the N position encoding information of the N second text unit Tokens are consecutive in sequence; Wherein, there are K groups of text unit Tokens with the same position encoding information among the N first text unit Tokens and the N second text unit Tokens; and, the text position indicated by the position encoding information of the i-th text unit Token among the N first text unit Tokens is located before the text position indicated by the position encoding information of the i-th text unit Token among the N second text unit Tokens, i is a positive integer and i is less than or equal to N.

7. The method according to claim 3, wherein The obtaining the second predicted text sequence based on the N first text unit Tokens includes: Obtain the t×N text unit Tokens associated with the t-th time step and the time steps before the t-th time step, and the t×N text unit Tokens include the N first text unit Tokens; Select M reference text unit Tokens corresponding to M position encoding information from the t×N text unit Tokens; the M position encoding information includes the last position encoding information and each position encoding information whose indicated text position is before the text position indicated by the last position encoding information, and the last position encoding information is the position encoding information with the text position indicated being the last among the N position encoding information of the N first text unit Tokens, and M is a positive integer; Construct the second predicted text sequence using the M reference text unit Tokens corresponding to the M position encoding information.

8. The method according to claim 7, wherein Any one of the M position encoding information is target position encoding information, and each of the t×N text units Token has its own prediction confidence; selecting M reference text units corresponding to the M position encoding information from the t×N text units Token includes: Obtaining one or more text units Token with the target position encoding information from the t×N text units Token; Taking the text unit Token with the highest prediction confidence among the one or more text units Token as the reference text unit Token corresponding to the target position encoding information; or, Taking the text unit Token with the most occurrences among the one or more text units Token as the reference text unit Token corresponding to the target position encoding information.

9. The method according to claim 3, characterized in that, Performing text prediction processing using the first reference feature sequence to generate N first text units Token associated with the t-th time step includes: Obtaining a trained language model; the trained language model includes a shared network and N prediction networks; Invoking the shared network to perform feature learning on the first reference feature sequence to generate a shared sequence feature of the first reference feature sequence; Invoking each prediction network to perform text prediction processing using the shared sequence feature to generate the N first text units Token associated with the t-th time step; Wherein, one prediction network is used to generate one of the N first text units Token associated with the t-th time step.

10. The method according to claim 1, characterized in that, The target prediction text is generated by invoking the trained language model for text prediction processing; the method further includes: Obtaining a language model and sample prompt information; Invoking the language model to perform text prediction processing based on the sample prompt information according to the sample time steps to generate N sample text units Token respectively associated with one or more of the sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text units Token respectively associated with each sample time step; Generating a text prediction loss of the language model for the predicted sample text units Token based on the N sample text units Token respectively associated with one or more of the sample time steps; Using the text prediction loss to correct the model parameters of the language model to obtain the trained language model.

11. The method according to claim 10, wherein The N sample text units Token respectively associated with one or more of the sample time steps are used to generate a sample reply text corresponding to the sample prompt information, and each of the N sample text units Token associated with one or more of the sample time steps has its own position encoding information, and the position encoding information of any one of the sample text units Token is used to indicate the text position of any one of the sample text units Token in the sample reply text when the sample reply text uses any one of the sample text units Token; Among them, the sample prompt information has label text, the label text is a reference for the sample response text, the label text contains L label text units Token, and the L label text units Token all have position encoding information. The position encoding information of any one of the label text units Token is used to indicate the text position of any one of the label text units Token in the label text, and L is a positive integer; Any one of the sample text units Token associated with any one of the sample time steps is the same as the label text unit Token to which the position encoding information of any one of the sample text units Token belongs among the L label text units Token. Any one of the sample text units Token has a prediction confidence at any one of the sample time steps.

12. The method according to claim 11, wherein Generating the text prediction loss of the language model for the predicted sample text units Token based on one or more of the sample time steps each associated with N sample text units Token includes: Based on the prediction confidences of the N sample text units Token respectively associated with each of the sample time steps, generating an intermediate prediction loss of the language model at each of the sample time steps; Performing a summation process on the intermediate prediction losses of the language model at each of the sample time steps to generate the text prediction loss.

13. The method according to claim 12, characterized in that, Any one of one or more of the sample time steps is the s-th sample time step, where s is a positive integer; generating the intermediate prediction loss of the language model at each of the sample time steps based on the prediction confidences of the N sample text units Token respectively associated with each of the sample time steps includes: Based on the prediction confidences of the N sample text units Token respectively associated with the s-th sample time step, generating a unit prediction loss corresponding to each of the sample text units Token associated with the s-th sample time step; Performing a summation process on the N unit prediction losses corresponding to the N sample text units Token associated with the s-th sample time step to generate the intermediate prediction loss of the language model at the s-th sample time step.

14. A text prediction device, characterized in that, The device includes: An acquisition module, configured to acquire prompt information for text prediction; A prediction module, configured to perform text prediction processing step by step based on the prompt information according to at least one time step to obtain a target prediction text; the text prediction processing corresponding to each time step is used to generate N text units Token respectively associated with each time step, and each text unit Token has its own position encoding information. There are K groups of text units Token with the same position encoding information between the N text units Token respectively associated with two adjacent time steps. N is an integer greater than 1, and K is a non-negative integer and K is less than N; A generation module, configured to generate a response text corresponding to the prompt information according to the target prediction text; the response text contains the text units Token generated by the text prediction processing corresponding to each time step.

15. A computer program product, characterized in that, The computer program product includes a computer program, which is stored in a computer-readable storage medium and is adapted to be read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-13.

16. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1-13.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is adapted to be loaded and executed by a processor to execute the steps of the method according to any one of claims 1-13.