Text prediction method and apparatus, product, and device

WO2026200222A1PCT designated stage Publication Date: 2026-10-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/072462
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-01-14
Publication Date
2026-10-01

Smart Images

  • Figure CN2026072462_01102026_PF_FP_ABST
    Figure CN2026072462_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A text prediction method, executed by a computer device. The method comprises: acquiring prompt information used for text prediction (S101); on the basis of the prompt information, performing text prediction processing step by step according to time steps to obtain a target predicted text, wherein the text prediction processing corresponding to each time step is used for generating N text units associated with said time step, each text unit has its own position encoding information, if there are a plurality of time steps, K groups of text units having the same position encoding information exist between N text units respectively associated with two adjacent time steps, N is an integer greater than 1, K is a non-negative integer, and K is less than N (S102); and on the basis of the target predicted text, generating a reply text corresponding to the prompt information, wherein the reply text comprises text units generated by means of text prediction processing corresponding to each time step (S103).
Need to check novelty before this filing date? Find Prior Art

Description

Text prediction methods, devices, products and equipment

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on March 24, 2025, with application number 202510352580.9 and entitled “Text Prediction Method, Apparatus, Product and Equipment”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence, and more particularly to a text prediction method, apparatus, product, and device. Background Technology

[0004] With the continuous development of AI (Artificial Intelligence), language models have made significant progress in the field of natural language processing. By learning contextual relationships from large-scale text corpora, language models have demonstrated powerful language generation and understanding capabilities. After a user inputs a prompt word, the language model can predict and generate corresponding words one by one, ultimately forming the response content for that prompt word. However, when the data volume of the text to be processed (such as prompt words) is very large, the efficiency of generating prompt word responses using this method of generating individual words becomes extremely low. Therefore, how to improve the efficiency of generating prompt word responses is a hot topic. Summary of the Invention

[0005] This application provides a text prediction method, apparatus, product, and device.

[0006] This application provides a text prediction method, executed by a computer device, the method comprising:

[0007] Obtain prompts for text prediction;

[0008] Based on the prompt information, text prediction processing is performed step by step according to time steps to obtain the target predicted text; the text prediction processing corresponding to each time step is used to generate N text units associated with each time step. Each text unit has its own position encoding information. If there are multiple time steps, there are K groups of text units with the same position encoding information among the N text units associated with two adjacent time steps. N is an integer greater than 1, K is a non-negative integer, and K is less than N.

[0009] The response text is generated based on the target predicted text, and contains text units generated by the text prediction processing at each time step.

[0010] This application provides a text prediction device, which includes:

[0011] The acquisition module is used to acquire prompt information for text prediction.

[0012] The prediction module is used to perform text prediction processing step by step according to the prompt information to obtain the target predicted text. The text prediction processing corresponding to each time step is used to generate N text units associated with each time step. Each text unit has its own position encoding information. If there are multiple time steps, there are K groups of text units with the same position encoding information among the N text units associated with two adjacent time steps. N is an integer greater than 1, K is a non-negative integer, and K is less than N.

[0013] The generation module is used to generate response text corresponding to the prompt information based on the target predicted text; the response text contains text units generated by the text prediction processing corresponding to each time step.

[0014] This application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the method of this application.

[0015] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method described in the above-mentioned aspect.

[0016] This application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in the various alternative embodiments described above.

[0017] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.

[0019] Figure 1 is a schematic diagram of the network architecture of a text generation network provided in an embodiment of this application;

[0020] Figure 2 is a schematic diagram of a scenario for generating reply text corresponding to prompt information according to an embodiment of this application;

[0021] Figure 3 is a flowchart illustrating a text prediction method provided in an embodiment of this application;

[0022] Figure 4 is a schematic diagram illustrating the principle of generating N first text units according to an embodiment of this application;

[0023] Figure 5 is a schematic diagram illustrating the principle of constructing a first reference feature sequence according to an embodiment of this application;

[0024] Figure 6 is a schematic diagram of a scenario for text prediction processing through prompt information provided in an embodiment of this application;

[0025] Figure 7 is a schematic diagram of a process for training a language model according to an embodiment of this application;

[0026] Figure 8 is a schematic diagram illustrating the principle of training a language model according to an embodiment of this application;

[0027] Figure 9 is a schematic diagram of the structure of a text prediction device provided in an embodiment of this application;

[0028] Figure 10 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] All data collected in this application (such as prompts, sample prompts, and other related data) are collected with the consent and authorization of the data owner (such as a user, organization, or enterprise), and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions.

[0031] Here, the relevant technical concepts involved in this application are explained:

[0032] Token: The smallest unit for text prediction, which can be referred to as a text unit in this application. One text unit is one token. A token can be a word, a phrase, or a character, etc.

[0033] Large Language Model (LLM) is a core deep learning technique in the field of AI (Artificial Intelligence). It is a deep learning model trained on a large amount of text data, capable of generating and understanding natural language. In this application, the large language model may be referred to as a language model.

[0034] Time step: A time step represents a prediction action of the large language model for a token. That is, a time step can be defined as an independent prediction step when the large language model generates text. This prediction step is used to predict and output a token.

[0035] Please refer to Figure 1, which is a schematic diagram of the network architecture of a text generation network provided in an embodiment of this application. As shown in Figure 1, the network architecture may include a server 200 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices; the number of terminal devices is not limited here. As shown in Figure 1, the multiple terminal devices may specifically include terminal device 1, terminal device 2, terminal device 3, ..., terminal device n. As shown in Figure 1, terminal device 1, terminal device 2, terminal device 3, ..., terminal device n can all be connected to the server 200 via the network, so that each terminal device can interact with the server 200 through the network connection.

[0036] As shown in Figure 1, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Terminal devices can be smart terminals such as smartphones, tablets, laptops, desktop computers, smart TVs, in-vehicle terminals, and smart home devices. The following description uses the communication between terminal device 1 and server 200 as an example to illustrate the specific implementation of this application.

[0037] Terminal device 1 may include a client (such as a question-and-answer client), and server 200 may be the backend server of this client. Users can input prompt information into the client; this prompt information is prompt text, such as a question. After receiving the prompt text, the client can send a reply request to server 200, which may carry the prompt text. The reply request is a request sent by the client to the server after receiving the prompt information input by the user; this request carries the prompt information and requests the server to generate a reply text corresponding to the prompt information. After receiving the reply request, server 200 can extract the prompt text from the reply request and generate the corresponding reply text, which is the text used to reply to the prompt text, such as an answer to the question posed in the prompt text. Server 200 can return the reply text to the client, allowing the client to display the reply text on its interface for user review. The process by which server 200 generates the reply text corresponding to the prompt text can be found in the descriptions of the following embodiments.

[0038] Please refer to Figure 2, which is a schematic diagram of a scenario for generating reply text corresponding to a prompt message according to an embodiment of this application. After obtaining the prompt text (i.e., the prompt message), the server 200 in Figure 1 can perform feature transformation on the prompt text, converting it into a feature sequence, which can be called the prompt feature sequence. For example, the server 200 can perform word segmentation on the prompt text to obtain a prompt text sequence, which is a sequence composed of the words obtained from the word segmentation of the prompt text. The server 200 can perform feature learning (such as feature embedding) on ​​each word in the prompt text sequence to obtain the feature vector of each word. A feature vector can be an embedding (embedded feature). Thus, the feature vectors of each word can constitute the prompt feature sequence, which contains the embeddings of each word in the prompt text. An embedding can be an element in the prompt feature sequence.

[0039] Server 200 can perform text prediction processing step by step according to the prompt feature sequence, thereby generating the target predicted text. This step-by-step text prediction processing can include one or more (here, multiple) text prediction processes corresponding to each time step. Each time step's text prediction process is used to generate multiple text units (which can be called tokens) associated with that time step. For example, through text prediction processes corresponding to multiple time steps (including time steps 1 to 6), multiple text units associated with time step 1, time step 2, time step 3, time step 4, time step 5, and time step 6 can be generated. The target predicted text can be obtained from the multiple text units associated with each time step. The specific process for obtaining the target predicted text can be further described in the relevant embodiment corresponding to Figure 3 below. Each text unit can have its own positional encoding information, which can be used to indicate which word the corresponding text unit predicts for the final text to be generated (i.e., the word predicted for which text position in the final text to be generated). Among multiple text units (such as N text units, where N is a positive integer) associated with each adjacent time step, there can be K groups of text units with the same positional encoding information, where K is a non-negative integer and K is less than N.

[0040] Therefore, server 200 can obtain the response text corresponding to the prompt information from the aforementioned target predicted text. For example, useless text units in the target predicted text can be removed to obtain the response text. Through the above process, server 200 can perform text prediction processing on the prompt text according to time steps, and finally obtain the response text corresponding to the prompt text.

[0041] Using the method provided in this application, multiple text units associated with each time step can be generated at each time step of text prediction processing based on prompt information. Therefore, by using the multiple text units generated at each time step, the corresponding reply text can be quickly obtained, thus improving the efficiency of text prediction processing and reducing the latency of sending a text reply to the user. Furthermore, there can be K groups of text units with the same positional encoding information among the multiple text units associated with adjacent time steps, where K is an adjustable value within the range [0, N). Different values ​​of K can bring different generation effects to the reply text. Therefore, by adjusting the value of K, this application can also meet various business requirements regarding the latency or accuracy of text prediction. A detailed explanation of this effect can be found in the relevant description of step S103 in the embodiment corresponding to Figure 3 below.

[0042] Please refer to Figure 3, which is a flowchart illustrating a text prediction method provided in an embodiment of this application. The executing entity in this embodiment can be a text generation device (hereinafter referred to as the generation device). The generation device can be a computer device or a cluster of multiple computer devices. The computer device can be a server or other devices; this application does not impose any limitations on this. As shown in Figure 3, the method may include:

[0043] Step S101: Obtain prompt information for text prediction.

[0044] Specifically, the generating device can obtain prompt information, which can be a prompt word for a large language model (such as the trained language model described below in this application). This prompt information can be sent to the generating device by a client, meaning the generating device can obtain the prompt information sent by the client. The generating device can be a backend device (such as a backend server) of the client, and the client can be any client that supports question answering, such as a question answering client.

[0045] The prompt information can be of any type used to guide the large language model in language understanding. For example, it can be text (such as prompt text) and / or image (such as prompt image), etc. The prompt information can be unimodal (such as prompt text or prompt image) or multimodal (such as containing both prompt text and prompt image), depending on the specific application scenario. For instance, the prompt text can be any text entered by the user on the client side, such as a question, a word, or any sentence entered by the user on the client side; the prompt image can be any image uploaded or imported by the user on the client side, such as an image of a person, an image of a plant, etc.

[0046] Step S102: Based on the prompt information, perform text prediction processing step by step according to the time step to obtain the target predicted text; the text prediction processing corresponding to each time step is used to generate N text units associated with each time step. Each text unit has its own position encoding information. If there are multiple time steps, there are K groups of text units with the same position encoding information among the N text units associated with two adjacent time steps. N is an integer greater than 1, K is a non-negative integer, and K is less than N.

[0047] Specifically, the generating device can perform text prediction processing progressively according to at least one time step using the above-mentioned prompt information to generate target predicted text. The target predicted text can be the predicted text that is ultimately used to generate the response text corresponding to the prompt information. The target predicted text can be the last predicted text sequence before the end of the text prediction processing process, as described below.

[0048] The text prediction process performed by the generation device according to time steps can include text prediction processing corresponding to one or more time steps (i.e., at least one time step). The text prediction processing corresponding to each time step can be used to generate N text units associated with each time step, where N is an integer greater than 1. That is, when performing text prediction processing at any time step, multiple text units (such as N text units) associated with that time step can be generated. The specific value of N can be flexibly set according to actual business needs. For example, N can be set to 3, 4, or 5.

[0049] Each text unit can have its own positional encoding information, and a text unit associated with a time step can have one positional encoding information. The positional encoding information of any text unit can be used to indicate the text position (e.g., the word number) of that text unit in the response text when it is used in a subsequent response text.

[0050] Between any two adjacent time steps (which can be any two adjacent time steps), there can be K groups of text units with the same positional encoding information among the N text units associated with each other. K is a non-negative integer, and K is less than N. The value of K can be set according to actual business needs. K can be any non-negative integer less than N, that is, K can be any integer in [0, N). In other words, there may or may not be text units with overlapping text positions among the N text units predicted by adjacent time steps.

[0051] The generation device can acquire a trained language model, which can be a large language model (such as an LLM model). The process of training this language model can be seen in the specific description of the embodiment corresponding to Figure 7 below. The generation device can call the trained language model to perform text prediction processing step-by-step according to the prompt information to generate the target predicted text. Specifically, the trained language model can be used to predict and generate one text unit in one time step. Specifically, the shared network in the trained language model first processes the prompt feature sequence F... h Feature learning is performed to obtain shared sequence features F. s Assuming the current time step is t, the shared sequence feature F will be used. sThe input is fed into N prediction networks, and the i-th prediction network f i The shared sequence features are processed to output feature F. pi Then, through the word embedding decoding matrix f u Mapped to the vocabulary dimension, the probability distribution P of candidate text units in the candidate text unit library at the i-th text position at time step t is obtained using the softmax function. fb The formula is P fb =softmax(f u (f i (F s The candidate text unit with the highest probability is selected as the text unit at that position. The trained language model can be used to predict and generate a text unit once per time step.

[0052] The word embedding decoding matrix is ​​a matrix in the trained language model, also known as the decoding matrix or decoding network. It is used to map the feature representations of the hidden layer (such as features obtained by the output of the prediction network) to the dimension of the vocabulary (i.e., the lexicon). It belongs to the model parameters of the language model. Each prediction network can share this matrix to perform feature decoding processing, thereby obtaining the predicted text units.

[0053] If there exists a time step, the N text units associated with that time step can be predicted through a cue feature sequence, which is generated by feature transformation of the cue information.

[0054] If multiple time steps exist, these time steps are sequential. The text prediction processing for any given time step can be based on the text prediction processing for previous time steps. That is, the N text units associated with any given time step can be predicted using the representation features of the text units generated from the time steps preceding that time step and the aforementioned prompt feature sequence. The text prediction processing for each time step can be performed sequentially. A text unit can be a predicted token (the smallest unit for text prediction), such as a predicted character or phrase.

[0055] The representation feature of any text unit is the feature used to decode that text unit. This feature can be generated and output by the prediction network in the language model described below.

[0056] In one implementation, the process of performing feature transformation on the prompt information to generate a prompt feature sequence may include: performing feature encoding (i.e. feature transformation) on the prompt information through an information encoder adapted to the information type of the prompt information to obtain the prompt feature sequence of the prompt information.

[0057] For example, if the prompt information is a prompt text, a text encoder can be used to perform feature encoding on each word in the prompt text to generate a feature vector (embedding) for each word. Thus, the prompt feature sequence of the prompt text can be constructed using the feature vectors of each word.

[0058] Specifically, if the prompt message is a prompt text, the text encoder segments the prompt text to obtain a word sequence W = [w1, w2, ..., w n ], then for each word w i Feature encoding is performed through the embedding layer E to obtain the feature vector e. i =E(w i The final prompt text's prompt feature sequence F h_text =[e1,e2,…,e n If the prompt message is a prompt image, the image encoder first performs a convolution operation C on the image, converting image I into a feature map F. map =C(I), and then obtain the feature vector e through pooling operation P and fully connected layer F. image =F(P(F) map The cue feature sequence F of the cue image is shown. h_image =[e image If the prompt information includes both prompt text and prompt image, the prompt feature sequence of the prompt text and the prompt feature sequence of the prompt image are concatenated to obtain the prompt feature sequence F. h =concat(F h_text ,F h_image ), where concat represents the concatenation operation.

[0059] For example, if the prompt information is a prompt image, and there are one or more prompt images, then each prompt image can be processed by feature encoding to generate a feature vector (embedding) for each prompt image. Thus, the feature vector of each prompt image can also form a corresponding prompt feature sequence.

[0060] For example, if the prompt information includes both the prompt text and the prompt image, the feature sequence obtained by feature transformation of the prompt text can be concatenated with the feature sequence obtained by feature transformation of the prompt image to obtain the prompt feature sequence. This prompt feature sequence can contain the feature vector of each word in the prompt text and the feature vector of each prompt image.

[0061] The aforementioned information encoder (such as a text encoder or an image encoder) can be a pre-trained model capable of feature encoding (i.e. feature transformation) of information of a suitable information type, or the information encoder can be part of a trained language model and trained together with the trained language model.

[0062] If the information encoder is a model outside the trained language model, the above-mentioned prompt feature sequence can be used as input to the trained language model; if the information encoder is part of the trained language model, the above-mentioned prompt information can be used as input to the trained language model.

[0063] In one implementation, any one of the at least one time step can be referred to as the t-th time step, where t is a positive integer. That is, the t-th time step can be any time step in the text prediction process. The following description uses the text prediction process corresponding to the t-th time step as an example. It can be understood that the principle of text prediction processing for each time step is the same, but the text prediction processes for each time step can be performed sequentially according to the order of the time steps.

[0064] The generating device can acquire a first predicted feature sequence, which can be composed of the representation features of text units in a first predicted text sequence. That is, the first predicted feature sequence can contain the representation features of each text unit in the first predicted text sequence. The first predicted text sequence can be composed of text units predicted by the prompt feature sequence at time steps prior to the t-th time step. If t equals 1, there are no other time steps before the t-th time step, and the first predicted text sequence can be empty, meaning it may not exist, and therefore the first predicted feature sequence also may not exist. This is because there are no text units predicted by the prompt feature sequence before the t-th time step. In other words, when t equals 1, there may not be a first predicted feature sequence; instead, there may be N text units associated with the first time step predicted by the prompt feature sequence. If t is greater than 1, the principle for acquiring the first predicted text sequence is the same as the principle for acquiring the second predicted text sequence described below. For details, please refer to the principle for acquiring the second predicted text sequence described below.

[0065] The generating device can construct a reference feature sequence corresponding to the t-th time step based on the first predicted feature sequence and the aforementioned prompt feature sequence. This reference feature sequence corresponding to the t-th time step can be referred to as the first reference feature sequence. For example, the first reference feature sequence can be a feature sequence obtained by concatenating the first predicted feature sequence and the prompt feature sequence; that is, the first predicted feature sequence and the prompt feature sequence can be concatenated (the concatenation order can be a preset order) to obtain the first reference feature sequence. The first reference feature sequence can be the input at the t-th time step, meaning that at the t-th time step, text prediction can be performed using the first reference feature sequence, as described below.

[0066] The generating device can call the trained language model to perform text prediction processing based on the first reference feature sequence to generate N text units associated at the t-th time step. The N text units associated at the t-th time step can be referred to as the N first text units.

[0067] The trained language model of this application can include a shared network and N prediction networks. The shared network can also be referred to as the backbone network of the trained language model, and the N prediction networks can share the backbone network. The generation device can input the aforementioned first reference feature sequence into the shared network to call the shared network to perform feature learning (which can be embedding) on ​​the first reference feature sequence to generate shared sequence features of the first reference feature sequence. These shared sequence features are the features learned by the shared network from the first reference feature sequence (belonging to implicit feature representation). The backbone network can be a complex Transformer network, and each prediction network can be a lightweight Transformer layer. The model parameters of the backbone network can be much greater than the model parameters of each prediction network. Therefore, from the perspective of the language model as a whole, the multiple prediction networks introduced in this application hardly change the number of model parameters of the language model, and the model parameters of each prediction network are a very small part of the model parameters of the entire language model.

[0068] The generation device can invoke each prediction network in the trained language model to perform text prediction processing using the shared sequence features, thereby generating N first text units associated with the t-th time step. One prediction network can be used to perform text prediction processing using the shared sequence features to generate one first text unit associated with the t-th time step. That is, the number of prediction networks can be equal to N. In actual business needs, the number of text units to be generated in a time step determines the number of prediction networks set in the language model. N prediction networks can be used to generate N text units associated with that time step in parallel and quickly.

[0069] Please refer to Figure 4, which is a schematic diagram illustrating the principle of generating N first text units according to an embodiment of this application. As shown in Figure 4, the trained language model may include a shared network and N prediction networks, where N can be equal to 4, i.e., there are 4 prediction networks, including prediction network 1 to prediction network 4. The generation device can input a first reference feature sequence into the shared network to call the shared network to perform feature learning on the first reference feature sequence, generating a shared sequence feature (an embedded feature) of the first reference feature sequence. This shared sequence feature can be input into each prediction network, and each prediction network can use the shared sequence feature to perform text prediction processing to generate corresponding predicted text features. These predicted text features are the features output by the prediction network using the shared sequence feature for text prediction processing. For example, prediction network 1 can generate predicted text feature 1, prediction network 2 can generate predicted text feature 2, prediction network 3 can generate predicted text feature 3, and prediction network 4 can generate predicted text feature 4.

[0070] Text unit 1 can be mapped from the predicted text feature 1, text unit 2 can be mapped from the predicted text feature 2, text unit 3 can be mapped from the predicted text feature 3, and text unit 4 can be mapped from the predicted text feature 4. The N first text units associated at the t-th time step can include text unit 1, text unit 2, text unit 3, and text unit 4. The principle of decoding the predicted text features obtained by the prediction network to obtain the corresponding text units can be found in the principle described in formula (1) below.

[0071] In this context, predicted text feature 1 can be a representation feature of text unit 1, predicted text feature 2 can be a representation feature of text unit 2, predicted text feature 3 can be a representation feature of text unit 3, and predicted text feature 4 can be a representation feature of text unit 4. That is, the representation feature of any text unit is the feature generated by the prediction network for decoding that text unit, and this feature can be an embedding feature.

[0072] The generating device can obtain a second predicted text sequence from the aforementioned N first text units, and can obtain the target predicted text from the second predicted text sequence. The properties of the second predicted text sequence are the same as those of the first predicted text sequence.

[0073] If the second predicted text sequence contains text units of the predicted end type (i.e., end symbols), it indicates that the text prediction process has ended and there is no need to perform text prediction processing for the next time step. The second predicted text sequence can be directly used as the target predicted text.

[0074] If the second predicted text sequence does not contain text units of the ending type, it indicates that the text prediction process is not yet complete and text prediction processing for the next time step is required. Therefore, the generation device can continue to construct a reference feature sequence for the (t+1)th time step based on the representation features of the text units in the second predicted text sequence and the aforementioned prompt feature sequence. For example, the generation device can construct the second predicted feature sequence based on the representation features of the text units in the second predicted text sequence; that is, the second predicted feature sequence is composed of the representation features of each text unit in the second predicted text sequence. The reference feature sequence for the (t+1)th time step can be called the second reference feature sequence. Similarly to the first reference feature sequence, the second reference feature sequence can also be obtained by concatenating the second predicted feature sequence and the prompt feature sequence. The generation device can call the trained language model to perform text prediction processing through the second reference feature sequence to generate N text units associated with the (t+1)th time step. The N text units associated with the (t+1)th time step can be called N second text units. The principle of generating N second text units by performing text prediction processing through the second reference feature sequence is the same as the principle of generating N first text units by performing text prediction processing through the first reference feature sequence.

[0075] Furthermore, the generating device can obtain a third predicted text sequence through the N second text units, and continue to generate the target predicted text through the third predicted text sequence. The principle of generating the target predicted text through the third predicted text sequence is the same as that of generating the target predicted text through the second predicted text sequence described above. It mainly involves determining whether the currently generated predicted text sequence contains text units of the end type, i.e., determining whether the text prediction process has ended. If not, it continues with the text prediction process corresponding to the next time step; if it has ended, the current predicted text sequence can be used as the final generated target predicted text.

[0076] Each of the aforementioned N first text units and N second text units can have its own positional encoding information; that is, each text unit can have its own positional encoding information, and each text unit can have one positional encoding information. The positional encoding information of any text unit can be used to indicate the text position of that text unit in the subsequent response text when that text unit is used. In other words, the positional encoding information of a text unit is used to indicate which word that text unit is predicted for in the response text. Simply put, the positional encoding information of a text unit is used to indicate the text position of that text unit, which can be understood as the position number of that text unit, such as the position number indicating which word that text unit is predicted for in the response text.

[0077] The N text positions indicated by the N positional encoding information of the N first text units are sequentially consecutive, meaning they are sequentially adjacent. Similarly, the N text positions indicated by the N positional encoding information of the N second text units are also sequentially consecutive, meaning they are also sequentially adjacent. In other words, the N text positions indicated by the N positional encoding information of the N text units generated at any given time step are all sequentially consecutive and adjacent; that is, the N text units generated at any given time step can be N words generated at that given time step that are sequentially consecutive.

[0078] In one embodiment, the aforementioned N first text units and N second text units may each have their own positional encoding information, and there may be K groups of text units with the same positional encoding information among the N first text units and N second text units.

[0079] And the text position indicated by the position encoding information of the i-th text unit (sorted according to the order of text positions) in the N first text units is located before the text position indicated by the position encoding information of the i-th text unit in the N second text units, where i is a positive integer and i is less than or equal to N.

[0080] In other words, the overall text position of the N text units associated with the previous time step must be before the overall text position of the N text units associated with the next time step. For example, the overall text position of the N first text units associated with the t-th time step must be before the overall text position of the N second text units associated with the (t+1)-th time step. That is, the next time step predicts and generates at least one more (or more) new text units at the following text positions compared to the previous time step.

[0081] For example, the aforementioned N first text units may include the first text unit w1 at the predicted first text position, the first text unit w2 at the predicted second text position, the first text unit w3 at the predicted third text position, and the first text unit w4 at the predicted fourth text position; and the aforementioned N second text units may include the second text unit w5 at the predicted third text position, the second text unit w6 at the predicted fourth text position, the second text unit w7 at the predicted fifth text position, and the second text unit w8 at the predicted sixth text position.

[0082] Therefore, there can be two groups of text units with the same positional encoding information among the N first text units and the N second text units, i.e., K equals 2. One group of text units with the same positional encoding information may include the first text unit w3 and the second text unit w5, which have the same positional encoding information and are both used to indicate the third text position; and another group of text units with the same positional encoding information may include the first text unit w4 and the second text unit w6, which have the same positional encoding information and are both used to indicate the fourth text position.

[0083] For another example, the aforementioned N first text units may include the first text unit w1 at the predicted first text position, the first text unit w2 at the predicted second text position, and the first text unit w3 at the predicted third text position; and the aforementioned N second text units may include the second text unit w4 at the predicted fourth text position, the second text unit w5 at the predicted fifth text position, and the second text unit w6 at the predicted sixth text position.

[0084] Therefore, there can be 0 groups of text units with the same positional encoding information among the N first text units and the N second text units, i.e., K equals 0. In this case, there are no text units with the same positional encoding information among the N first text units and the N second text units.

[0085] Therefore, the process of obtaining the second predicted text sequence through the aforementioned N first text units can include: the generating device can obtain t×N text units associated with the t-th time step and the time steps preceding the t-th time step (if any), that is, there are currently t time steps in total, and each time step has N associated text units, so t time steps are associated with t×N text units. These t×N text units can include the N first text units associated with the t-th time step.

[0086] The generating device can select M reference text units corresponding to positional encoding information from the t×N text units. These M positional encoding information may include the last positional encoding information and the positional encoding information of each text unit whose indicated text position precedes the text position indicated by the last positional encoding information. The last positional encoding information may be the positional encoding information of the last text unit among the N positional encoding information of the first text units. M is a positive integer.

[0087] For example, according to the order of text positions, the aforementioned N first text units may sequentially include the first text unit at the predicted second text position, the first text unit at the third text position, the first text unit at the fourth text position, and the first text unit at the fifth text position. Then, the position encoding information indicating the fifth text position can be the last position encoding information. The M position encoding information may include this last position encoding information, the position encoding information indicating the fourth text position, the position encoding information indicating the third text position, the position encoding information indicating the second text position, and the position encoding information indicating the first text position. In this case, the N text units associated with the previous time step of the t-th time step may include the text unit at the predicted first text position, the text unit at the second text position, the text unit at the third text position, and the text unit at the fourth text position.

[0088] Any one of the M location encoding information mentioned above can be called the target location encoding information. Since the principle of obtaining the reference text unit corresponding to each location encoding information is the same, the following description will take obtaining the reference text unit corresponding to the target location encoding information as an example. One location encoding information can correspond to one reference text unit.

[0089] The generating device can extract one or more text units with target location encoding information from the aforementioned t×N text units. The generating device can do this by constructing a dictionary (dict) with location encoding information as keys and a list of corresponding text units as values. pc The t×N text units are classified and stored according to their positional encoding information. The specific construction process is as follows: Initialize an empty dictionary dict. pc Iterate through t×N text units. For each text unit u and its position encoding information pc, if pc is not in the dictionary... pc In the middle, then dict pc [pc] = [u]; if pc is already in the dictionary. pc In the middle, add u to the dictionary. pc The list [pc], i.e., dict pc [pc].append(u). When it is necessary to obtain pc with target location encoding information. target When dealing with one or more text units, the corresponding value dict is retrieved directly from the dictionary. pc [pc target ], which refers to one or more text units that have target location encoding information.

[0090] If there are multiple time steps, and there are text units with overlapping text positions among the N text units associated with different time steps, then there can be multiple text units with target position encoding information; however, if there is only one time step, or there are no text units with overlapping text positions among the N text units associated with different time steps, then there can be only one text unit with target position encoding information.

[0091] In this context, each of the aforementioned t×N text units can have its own prediction confidence level. A text unit can have one prediction confidence level, which is the confidence level obtained when the text unit is predicted at the corresponding time step. The prediction confidence level of a text unit can also be the predicted probability of that text unit. A higher prediction confidence level indicates that the trained language model believes the corresponding text position contains that text unit, and vice versa.

[0092] Therefore, the generating device selects the text unit with the highest prediction confidence from one or more text units containing target location encoding information as the reference text unit corresponding to the target location encoding information. Alternatively, the generating device can select the text unit with the highest frequency (i.e., the number of occurrences) among these one or more text units as the reference text unit corresponding to the target location encoding information. The specific selection of the reference text unit corresponding to the target location encoding information can be flexibly configured according to the actual application scenario.

[0093] The generating device can use the M reference text units corresponding to the M location encoding information obtained above to construct a second predicted text sequence, which can include the M reference text units. The arrangement order of the M reference text units in the second predicted text sequence can be the order of the text positions indicated by the location encoding information corresponding to the M reference text units.

[0094] As can be seen from the above, the second predicted text sequence is the optimal text sequence verified by the text units generated in the previous time step (such as by the prediction confidence). The text sequence can be composed of the optimal text units at each of the currently predicted text positions. For example, a text unit with a higher prediction confidence can be considered the optimal text unit, or a text unit that appears more frequently at the same text position can be considered the optimal text unit at that text position.

[0095] Since N consecutive text units are generated at each time step, the order of the text positions generated by the N prediction networks can be set in this application. For example, the N prediction networks may include prediction network 1, prediction network 2, prediction network 3, and prediction network 4. For the four consecutive text positions for which text units need to be generated, prediction network 1 can be used to generate the text unit at the first text position, prediction network 2 can be used to generate the text unit at the second text position, prediction network 3 can be used to generate the text unit at the third text position, and prediction network 4 can be used to generate the text unit at the fourth text position.

[0096] Specifically, the N prediction networks can also have a set order. The i-th prediction network among the N prediction networks can be used to generate the text unit at the i-th position among the N consecutive text positions that need to be generated. i is a positive integer less than or equal to N.

[0097] The i-th prediction network can be represented as f i The aforementioned backbone network can be represented as Z. zg The aforementioned first reference feature sequence can be represented as x ck Through the backbone network and the i-th prediction network, the i-th text position can be generated from among the N consecutive text positions where text units need to be generated at time step t. This generates the prediction confidence for each candidate text unit in the candidate text unit library (i.e., vocabulary, which can contain a large number of candidate text units, i.e., a large number of words). The prediction confidence of each candidate text unit can collectively form a probability distribution. Finally, the candidate text unit with the highest prediction confidence in the candidate text unit library is selected as the text unit generated at time step t at the i-th text position. Therefore, the probability distribution P for generating candidate text units in the candidate text unit library at the i-th text position is... fb The principle can be illustrated by the following formula: P fb =softmax(f u (f i (Z zg (x ck )))) (1)

[0098] Among them, P fbThis can be included at the i-th text position, representing the predicted confidence score for each candidate text unit in the candidate text unit library. `softmax` is the normalization function. The normalization function processes the probability distribution obtained by mapping the output features of the prediction network through the word embedding decoding matrix, ensuring that the sum of the predicted confidence scores of all candidate text units is 1. This facilitates selecting the unit with the highest predicted confidence score from the candidate text unit library as the text unit at the corresponding position. u This can be the word embedding decoding matrix (or simply the decoding matrix, or the decoding network) in the trained language model. This decoding matrix maps the hidden layer's feature representations (such as features obtained from the prediction network's output) to the vocabulary (i.e., the lexicon). This decoding matrix is ​​part of the language model's parameters. Each prediction network can share this decoding matrix for feature decoding processing to obtain the predicted text units.

[0099] As can be understood from the above, when predicting text units at any time step, the N positional encoding information of the N consecutive text positions that need to be predicted are known in advance. Therefore, based on the prior availability of these N positional encoding information, text units corresponding to each of the N positional encoding information can be predicted and generated, thereby generating the N text units associated with that time step. The text unit corresponding to a generated positional encoding information can possess that positional encoding information, which is used to indicate the text position of the text unit in the final text to be generated, that is, indicating which word in the final text to be generated if the text unit is used in the final text to be generated (such as the response text below).

[0100] Please refer to Figure 5, which is a schematic diagram illustrating the principle of constructing a first reference feature sequence according to an embodiment of this application. As shown in Figure 5, the aforementioned t-th time step can be the 4th time step, and there can be 1st to 3rd time steps before the 4th time step. Assuming N equals 3 and K equals 2, the N text units generated by the text prediction processing corresponding to the 1st time step include text unit 1, text unit 2, and text unit 3; the N text units generated by the text prediction processing corresponding to the 2nd time step include text unit 11, text unit 22, and text unit 33; and the N text units generated by the text prediction processing corresponding to the 3rd time step include text unit 111, text unit 222, and text unit 333. That is, there can be two groups of text units with the same positional encoding information between any two adjacent time steps associated with the N text units.

[0101] Thus, text unit 2 and text unit 11 may have the same positional encoding information; text unit 3, text unit 22, and text unit 111 may have the same positional encoding information; and text unit 33 and text unit 222 may have the same positional encoding information.

[0102] Specifically, the positional encoding information for text unit 1 is positional encoding information 1, and text unit 1 is predicted from the first word; the positional encoding information for text units 2 and 11 is positional encoding information 2, and text units 2 and 11 are predicted from the second word; the positional encoding information for text units 3, 22, and 111 is positional encoding information 3, and text units 3, 22, and 111 are predicted from the third word; the positional encoding information for text units 33 and 222 is positional encoding information 4, and text units 33 and 222 are predicted from the fourth word; and the positional encoding information for text unit 333 is positional encoding information 5, and text unit 333 is predicted from the fifth word. This indicates that there are a total of 5 predicted words for text positions before the 4th time step, and these 5 text positions are the text positions of the first 5 words.

[0103] Therefore, if text unit 2 has the highest prediction confidence among text units 2 and 11, text unit 22 has the highest prediction confidence among text units 3, 22 and 111, and text unit 33 has the highest prediction confidence among text units 33 and 222, then the reference text unit corresponding to position coding information 1 can be text unit 1, the reference text unit corresponding to position coding information 2 can be text unit 2, the reference text unit corresponding to position coding information 3 can be text unit 22, the reference text unit corresponding to position coding information 4 can be text unit 33, and the reference text unit corresponding to position coding information 5 can be text unit 333.

[0104] The aforementioned first predicted text sequence can be composed of the reference text unit corresponding to position encoding information 1 (i.e., text unit 1), the reference text unit corresponding to position encoding information 2 (i.e., text unit 2), the reference text unit corresponding to position encoding information 3 (i.e., text unit 22), the reference text unit corresponding to position encoding information 4 (i.e., text unit 33), and the reference text unit corresponding to position encoding information 5 (i.e., text unit 333). Furthermore, by concatenating the first predicted feature sequence, composed of the representation features of each text unit in the first predicted text sequence, with the aforementioned prompt feature sequence, the input for the fourth time step, i.e., the aforementioned first reference feature sequence corresponding to the fourth time step, can be obtained.

[0105] At the fourth time step, text prediction processing can be performed using the first reference feature sequence to generate N text units associated with the fourth time step. Similarly, the N text units associated with the fourth time step can include text units predicted from the fourth word, text units predicted from the fifth word, and text units predicted from the sixth word.

[0106] Step S103: Generate a response text corresponding to the prompt information based on the target predicted text; the response text contains text units generated by the text prediction processing corresponding to each time step.

[0107] Specifically, if there are at least one other text unit following the aforementioned ending text unit in the target predicted text, then the at least one text unit can be a zero-value text unit (i.e., a 0 element) added by the trained language model after the ending text unit during text prediction processing. This ending text unit is usually an indicator (i.e., an end marker) indicating the end of prediction, and it is usually not reflected in the final response text.

[0108] The generating device can first locate the position pos of the end-type text unit in the target predicted text, and then remove the elements in the target predicted text with an index greater than or equal to pos to obtain the response text R. The formula is R = target_text[:pos], where target_text represents the target predicted text and [:pos] represents the elements in the target predicted text with an index less than pos.

[0109] Therefore, the generating device can remove the text units of the ending type and the zero-value text units that follow the ending type from the target predicted text to obtain the response text. This response text can contain the text units generated by the text prediction processing corresponding to each time step described above. In this response text, there can be one text unit at a text position where a prediction is performed.

[0110] The response text can be obtained by sequentially combining all text units in the target predicted text, excluding terminators and zero-value text units, according to the order of their corresponding text positions. This combination can mean merging them together (e.g., linking them together). This response text is generated by the trained language model through text prediction processing to reply to the prompt information. For example, if the prompt information is a question posed by the user, the response text can be the answer to that question; or if the prompt information is a word entered by the user, the response text can be a text explaining and describing the meaning of that word and related events; and so on.

[0111] The generating device can return the above reply text to the client, allowing the client to output the reply text in the client interface for the user to view and use, such as supporting the user to forward, comment, provide feedback, share, etc.

[0112] In one implementation scenario, a larger value for K results in more groups of text units with overlapping text positions among the N text units generated at adjacent time steps. This allows the language model to learn more features related to the text positions before and after, thus improving the accuracy of text prediction. Conversely, a smaller value for K results in fewer groups of text units with overlapping text positions among the N text units generated at adjacent time steps. This allows the language model to generate more text units at different text positions in fewer time steps, reducing the total number of steps required for text prediction and thus improving its efficiency. Therefore, in practical business scenarios, the value of K can be set based on actual business needs, considering both the accuracy and efficiency of text prediction. The effect of K can also be seen in the following language model training process: a larger value for K results in more groups of text units with the same label among the N sample text units generated at adjacent sample time steps.

[0113] Where K can be any integer between [0, N). A larger K indicates that the language model learns more correlation features between preceding and following text positions, resulting in fewer text units generated by the language model at a new text position in a given time step. This leads to a slower thinking speed for the language model. Therefore, a strategy with a larger K value can be called a slow-thinking strategy. In a slow-thinking strategy, the accuracy of the text units predicted by the language model for each text position can be higher because text units at the same text position can be generated at multiple time steps, allowing for progressive optimization at different time steps. Conversely, a smaller K indicates that the language model learns fewer correlation features between preceding and following text positions, resulting in more text units generated by the language model at a new text position in a given time step. This leads to a faster thinking speed for the language model. Therefore, a strategy with a smaller K value can be called a fast-thinking strategy.

[0114] In one implementation, when latency requirements outweigh accuracy requirements, a smaller K value results in higher text prediction efficiency. For example, K can be taken within a first value range (e.g., [0, N / 2)). This means K can be a value less than N / 2, which can be understood as employing the aforementioned fast-thinking strategy. Conversely, when accuracy requirements outweigh latency requirements, a larger K value results in higher text prediction accuracy. For example, K can be taken within a second value range (e.g., (N / 2, N)). This means K can be a value greater than N / 2, which can be understood as employing the aforementioned slow-thinking strategy. Values ​​within the first value range can generally be smaller than values ​​within the second value range.

[0115] The specific value of K (i.e., whether to use a slow-thinking or fast-thinking strategy) can be flexibly adjusted and changed according to the business needs of the actual business scenario. For example, the specific value of K can be set based on the relevant business attributes of text prediction, such as those related to the latency or accuracy of text prediction. For instance, when the network quality between the generating device and the client is poor, a fast-thinking strategy can be used for text prediction to respond to users faster and avoid further latency; conversely, in more specialized technical question-and-answer scenarios where the accuracy of the response is crucial, a slow-thinking strategy can be used for text prediction. And so on.

[0116] Please refer to Figure 6, which is a schematic diagram of a scenario for text prediction processing using prompt information provided in an embodiment of this application. As shown in Figure 6, the prompt information can be prompt text (which can be segmented into words including the words "knowledge", "then", "is", "power", and the added terminator "Esep"). The shared network can contain a text encoder. The prompt text can be input into the shared network (which can be denoted as Shared Trunk) to call the shared network to perform feature transformation on the prompt text, obtaining the prompt feature sequence of the prompt text. Furthermore, the shared network can be called to perform feature learning on the prompt feature sequence (which can be done at the first time step), and the features generated by the shared network's feature learning (such as the shared sequence features mentioned above) can be input into each prediction network for text prediction processing. The prediction network can also be called an independent prediction head, which can be denoted as Independent Output Head. Here, there are four independent prediction heads, i.e., N is 4. These four independent prediction heads can include Independent Output Head1, Independent Output Head2, Independent Output Head3, and Independent Output Head4.

[0117] For example, if the above-mentioned slow-thinking strategy is adopted, then through the text prediction networks of each prediction network, text units at the text positions indicated by position encoding information T1 (which could be the position of the first word), T2 (which could be the position of the second word), T3 (which could be the position of the third word), and T4 (which could be the position of the fourth word) can be generated at the first time step; and text units at the text positions indicated by position encoding information T2 (which could be the position of the second word), T3 (which could be the position of the third word), and T4 (which could be the position of the fourth word) can be generated at the second time step. The text unit at the current position (which could be the position of the 3rd word), the text unit at the text position indicated by position encoding information T4 (which could be the position of the 4th word), and the text unit at the text position indicated by position encoding information T5 (which could be the position of the 5th word) can be generated at the 3rd time step. Furthermore, text units at the text positions indicated by position encoding information T3 (which could be the position of the 3rd word), T4 (which could be the position of the 4th word), T5 (which could be the position of the 5th word), and T6 (which could be the position of the 6th word) can be generated at the 3rd time step. This process continues, generating text units associated with more time steps. The text prediction process ends when a text unit of the final type is predicted.

[0118] In this slow-thinking strategy, K is set to the maximum value of 3 within the range [0, N). Between two adjacent time steps, three groups of text units associated with the same positional encoding information can be obtained. Through the text units generated at each time step, and following the principles described above, a highly accurate response text corresponding to the prompt text can ultimately be obtained.

[0119] Furthermore, if the aforementioned fast-thinking strategy is adopted, then through the text prediction networks of each prediction network, text units at the text positions indicated by position encoding information T1 (which could be the position of the first word), T2 (which could be the position of the second word), T3 (which could be the position of the third word), and T4 (which could be the position of the fourth word) can be generated at the first time step; and text units at the text positions indicated by position encoding information T5 (which could be the position of the fifth word) and T6 (which could be the position of the fifth word) can be generated at the second time step. The text unit could be at the position of the 6th word, the position indicated by position encoding information T7 (which could be the position of the 7th word), or the position indicated by position encoding information T8 (which could be the position of the 8th word). Furthermore, at the 3rd time step, text units can be generated at the positions indicated by position encoding information T9 (which could be the position of the 9th word), T10 (which could be the position of the 10th word), T11 (which could be the position of the 11th word), and T12 (which could be the position of the 12th word). This process continues, generating text units associated with more time steps. The text prediction process ends when a text unit of the final type is predicted.

[0120] In this fast-thinking strategy, K is set to the minimum value of 0 within the range [0, N). There are 0 sets of text units with the same positional encoding information between the 4 text units associated with each of two adjacent time steps; that is, there are no sets of text units with the same positional encoding information between the 4 text units associated with each of two adjacent time steps. Through the text units generated at each time step, and following the principles described above, the corresponding response text can be obtained efficiently and quickly.

[0121] This application proposes a multi-token prediction method that can simultaneously predict multiple future tokens at various time steps. Through a shared backbone network and N independent output head structures (i.e., prediction networks, also known as prediction heads), the language model can learn and capture more global semantic information from the context at each time step by referencing multiple tokens associated with previous time steps. This reduces the language model's over-reliance on historical generation results (such as tokens at previous text positions), thereby reducing the accumulated error in the text generation process and significantly improving the accuracy of text prediction for generating text units. Furthermore, since this application can predict multiple future tokens synchronously and in parallel at each time step, it can also greatly improve the speed of text prediction for generating text units, thus enhancing the efficiency of text prediction.

[0122] Based on this, in natural language generation tasks, the multi-token prediction method of this application learns and captures more global semantic information from the context, thus reducing semantic drift in text generation, improving the coherence and consistency of the generated text content, and enhancing the generalization ability of the language model for text prediction. Semantic drift is a phenomenon in natural language generation tasks where the generated text content gradually deviates semantically from the original topic or context.

[0123] For users, the method provided in this application allows for the generation of response text corresponding to user prompts with extremely low latency, enabling fast and efficient responses and improving user experience. In online service text generation tasks, the latency of this application is significantly optimized from 2.6 seconds to 1.37 seconds compared to using traditional language models (such as a language model that generates one token per time step), essentially halving the latency and demonstrating superior text generation performance, greatly improving the efficiency of text generation.

[0124] This application can obtain prompt information for text prediction; and can perform text prediction processing step by step according to the prompt information to obtain the target predicted text; the text prediction processing corresponding to each time step is used to generate N text units associated with each time step, each text unit having its own positional encoding information. If there are multiple time steps, there are K groups of text units with the same positional encoding information among the N text units associated with two adjacent time steps, where N is an integer greater than 1, K is a non-negative integer, and K is less than N; and, it can also generate a response text corresponding to the prompt information according to the target predicted text; the response text contains the text units generated by the text prediction processing corresponding to each time step. Therefore, it can be seen that the method proposed in this application can generate multiple text units at each time step during the text prediction process. Furthermore, there can be K groups of text units with the same positional coding information among the N text units associated with each of two adjacent time steps. K can take any non-negative integer less than N. The larger the value of K, the more text units with the same positional coding information can be among the text units associated with each of two adjacent time steps, resulting in more correlation features between time steps and higher accuracy in text prediction. Conversely, the smaller the value of K, the fewer text units with the same positional coding information can be among the text units associated with each of two adjacent time steps, resulting in more text units with new positional coding information generated at each time step and higher efficiency in text prediction. It is evident that by generating multiple text units at each time step and setting a certain number (e.g., K groups) of text units with the same positional coding information between adjacent time steps, the efficiency and accuracy of text prediction can be improved using the method provided in this application.

[0125] Please refer to Figure 7, which is a schematic diagram of a process for training a language model according to an embodiment of this application. The target predicted text mentioned above can be generated by calling the trained language model for text prediction processing. Therefore, this embodiment of the application specifically describes how to train the language model to obtain the trained language model. As shown in Figure 7, the process may include:

[0126] Step S201: Obtain the language model and sample prompt information.

[0127] Specifically, the generating device can acquire the language model and sample prompt information. The concept of this sample prompt information can be the same as the concept of prompt information mentioned above; this sample prompt information is used as a sample for training the language model.

[0128] The sample prompt message may have a tag text, which can be set for the sample prompt message and represents the desired sample response text to be generated based on it. The relationship between the sample response text and the sample prompt message is similar to the relationship between the response text and the prompt message described above; the sample response text can be the text used to respond to the sample prompt message.

[0129] Step S202: The language model is invoked to perform text prediction processing based on the sample prompt information according to the sample time step, generating N sample text units associated with one or more sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text units associated with each sample time step.

[0130] Specifically, the generation device can invoke a language model to perform text prediction processing based on the aforementioned sample prompts according to the sample time steps, in order to generate N sample text units associated with one or more sample time steps. The concept of this sample time step is similar to the concept of time step mentioned above; it is simply referred to as a sample time step here to distinguish it from the training process.

[0131] The principle of calling the language model to perform text prediction processing according to sample time steps based on sample prompts, in order to generate N sample text units associated with each of one or more sample time steps, is the same as the principle of calling the trained language model to perform text prediction processing according to time steps based on prompts, in order to generate N text units associated with each of one or more time steps. The text prediction process also ends when a text unit of the ending type is predicted. These one or more sample time steps refer to the various sample time steps performed before the text prediction process ends. The specific number of sample time steps performed before the text prediction process ends can be determined based on the actual text prediction results of the language model.

[0132] Similarly, the process of text prediction processing according to the sample time step can also include text prediction processing corresponding to one or more sample time steps. The text prediction processing corresponding to each sample time step is used to generate N sample text units associated with each sample time step.

[0133] The N sample text units associated with each of the one or more sample time steps can be used to generate the text prediction loss of the language model for the predicted sample text units. This text prediction loss can be used to reflect the prediction bias of the language model for the predicted sample text units at each sample time step. The larger the text prediction loss, the larger the prediction bias; the smaller the text prediction loss, the smaller the prediction bias.

[0134] Similarly, the N sample text units associated with each of the one or more sample time steps can be used to generate sample response text corresponding to the sample prompt information. The principle of generating the sample response text through the N sample text units associated with each sample time step is the same as the principle of generating the response text through the N text units associated with each time step, such as both including the principle of obtaining the second predicted text sequence mentioned above.

[0135] The aforementioned label text can serve as a reference for the sample response text; ideally, the sample response text should match the label text. In actual training, this sample response text may not be generated. It is introduced here to illustrate the positional encoding information of each sample text unit.

[0136] The concept of positional encoding information for sample text units can be similar to the concept of positional encoding information for text units described above. Each sample text unit associated with one or more sample time steps can have its own positional encoding information. For example, the positional encoding information of any sample text unit can also be used to indicate the text position of that sample text unit in the sample response text when that sample text unit is used. In other words, the positional encoding information of a sample text unit can be used to indicate which word that sample text unit is predicted to be. There can also be K sets of sample text units with the same (i.e., overlapping) positional encoding information between the N sample text units associated with two adjacent sample time steps.

[0137] The aforementioned tag text can contain L tag text units, each of which can have positional encoding information. The positional encoding information of any tag text unit can indicate its position within the tag text. For example, the positional encoding information of any tag text unit can indicate which word it is within the tag text. A tag text unit is the most ideal text unit to be predicted at the text position indicated by its positional encoding information. For instance, if the positional encoding information of a tag text unit indicates that it is the 3rd word in the tag text, then predicting and generating the 3rd word as the tag text unit is the ideal situation during text prediction processing. L is a positive integer, and its specific value can be determined based on the actual application scenario. L can be the total number of text units that need to be predicted, generated, and used in the final process (such as the text units constituting the sample response text).

[0138] Any sample text unit associated with any sample time step can have the same positional encoding information as the corresponding label text unit among the L label text units, and this sample text unit can have a prediction confidence at that sample time step. In other words, during training, when generating sample text units at each sample time step, the label text units with the same positional encoding information can be selected from the L label text units based on the positional encoding information of the sample text unit to be generated (i.e., the one that needs to be generated), and used as the currently predicted sample text unit.

[0139] When generating a sample text unit corresponding to any positional encoding information at a sample time step, the prediction confidence of each candidate text unit in the candidate text unit library (i.e., vocabulary) is obtained. Therefore, at this time, the candidate text unit in the candidate text unit library that is the same as the label text unit to which the positional encoding information belongs can be used as the predicted sample text unit. The prediction confidence of this sample text unit is not necessarily the highest among the prediction confidence of each candidate text unit in the candidate text unit library.

[0140] For example, there can be a total of 3 sample time steps. The N sample text units associated with the first sample time step can include: sample text unit y1 at the first text position (i.e., the first word predicted in the first sample time step as a whole), sample text unit y2 at the second text position (i.e., the second word predicted in the first sample time step as a whole), and sample text unit y3 at the third text position (i.e., the third word predicted in the first sample time step as a whole).

[0141] The N sample text units associated with the second sample time step may include: sample text unit y11 at the third text position (i.e., the third word predicted in the second sample time step as a whole), sample text unit y22 at the fourth text position (i.e., the fourth word predicted in the second sample time step as a whole), and sample text unit y33 at the fifth text position (i.e., the fifth word predicted in the second sample time step as a whole).

[0142] Similarly, the N sample text units associated with the third sample time step can include: sample text unit y111 at the fifth text position (i.e., the fifth word predicted in the third sample time step as a whole), sample text unit y222 at the sixth text position (i.e., the sixth word predicted in the third sample time step as a whole), and sample text unit y333 at the seventh text position (i.e., the seventh word predicted in the third sample time step as a whole).

[0143] In the example above, there can be one group of N sample text units associated with two adjacent sample time steps that have the same positional encoding information (i.e., the same text position), where K equals 1. For example, the two adjacent sample time steps can include the first sample time step and the second sample time step, as well as the second sample time step and the third sample time step.

[0144] If the sample text unit y222 is an end-of-text unit, then the tag text can include a total of 5 tag text units from the 1st to the 5th text position. These 5 tag text units can include: tag text unit b1 at the 1st text position, tag text unit b2 at the 2nd text position, tag text unit b3 at the 3rd text position, tag text unit b4 at the 4th text position, and tag text unit b5 at the 5th text position.

[0145] Therefore, the sample text unit y1 can be the label text unit b1, sample text unit y2 can be the label text unit b2, sample text unit y3 can be the label text unit b3, sample text unit y11 can be the label text unit b3, sample text unit y22 can be the label text unit b4, sample text unit y33 can be the label text unit b5, and sample text unit y111 can also be the label text unit b5. Since sample text unit y222 is a text unit of the end type, sample text unit y333 can be a zero-value text unit (i.e., a supplementary 0 element).

[0146] Each sample text unit associated with each sample time step has its own prediction confidence at the associated sample time step.

[0147] Step S203: Based on N sample text units associated with one or more sample time steps, generate the text prediction loss of the language model for the predicted sample text units.

[0148] Specifically, the generation device can generate the language model's overall text prediction loss for the predicted sample text units based on the N sample text units associated with each of the above sample time steps, as described below. The specific steps are as follows: First, based on the prediction confidence of the N sample text units associated with each sample time step, the intermediate prediction loss of the language model at each sample time step is generated. Assume there are a total of z sample time steps, and the prediction confidence of the N sample text units associated with the s-th sample time step is P... s1 ,P s2 ,…,P sN Then the intermediate prediction loss L at the s-th sample time step s for Then, the intermediate prediction losses of the language model at each sample time step are summed to generate the text prediction loss L. z The formula is

[0149] The generation device can generate the intermediate prediction loss of the language model at each sample time step by using the prediction confidence of the N sample text units associated with each sample time step. The language model can have an intermediate prediction loss at a sample time step, which reflects the prediction bias of the N sample text units associated with that sample time step.

[0150] Any one of the above-mentioned sample time steps can be referred to as the s-th sample time step, where s is a positive integer, meaning the s-th sample time step can be any of the above sample time steps. Since the principle of generating the intermediate prediction loss of the language model at each sample time step using the prediction confidence of the N sample text units associated with each sample time step is the same, this section will take the process of generating the intermediate prediction loss of the language model at the s-th sample time step using the prediction confidence of the N sample text units associated with the s-th sample time step as an example for specific explanation.

[0151] The generation device can generate the unit prediction loss corresponding to each sample text unit associated with the s-th sample time step by using the prediction confidence of each of the N sample text units associated with the s-th sample time step. Therefore, the N unit prediction losses corresponding to the N sample text units associated with the s-th sample time step can be summed (i.e., summed) to generate the intermediate prediction loss of the language model at the s-th sample time step.

[0152] In one implementation, based on a conditional independence assumption, the distribution of the N prediction confidences (i.e., prediction probabilities) of the N sample text units associated with the s-th sample time step can be decomposed and represented as the product of the prediction confidences of the N sample text units, as shown in the following formula:

[0153] Among them, P θ P can be the product of the N prediction confidences of the N sample text units associated with the s-th sample time step. i It can represent the prediction confidence of the i-th sample text unit among the N sample text units, where i is a positive integer less than or equal to N.

[0154] Therefore, the intermediate prediction loss L of the language model at the s-th sample time step s It can be:

[0155] In the above formula, log represents taking the logarithm, log(P) i ) represents the unit prediction loss corresponding to the i-th sample text unit above, L s This represents the intermediate prediction loss of the language model at the s-th sample time step.

[0156] The generation device can obtain the intermediate prediction loss of the language model at each sample time step (all sample time steps experienced at the end of the text prediction process) according to the above principle. The generation device can sum up the intermediate prediction losses of the language model at each sample time step (i.e., summation) to generate the overall text prediction loss of the language model.

[0157] Therefore, assuming there are a total of z sample time steps, where z is a positive integer, the final overall text prediction loss of the language model is L. z It can be:

[0158] Where j represents the j-th sample time step, j is a positive integer less than or equal to z, L z This represents the final text prediction loss of the language model. This text prediction loss reflects the overall prediction bias of the language model for the individual text units associated with each sample time step.

[0159] Step S204: Use text prediction loss to correct the model parameters of the language model to obtain the trained language model.

[0160] Specifically, the generation device can use the text prediction loss obtained above to correct the model parameters of the language model, and finally obtain the trained language model. The trained language model can then be applied to text generation in actual text prediction processing scenarios, such as generating the reply text corresponding to the above prompt information.

[0161] The generating device can employ the stochastic gradient descent (SGD) algorithm to utilize the text prediction loss L z Adjust the model parameters θ of the language model. The specific steps are: first, calculate the text prediction loss L. z gradient with respect to model parameters θ Then, update the model parameters based on the learning rate α, using the following formula: By iteratively updating the model parameters, the trained language model is finally obtained. This trained language model can then be applied to text generation in real-world text prediction and processing scenarios, such as generating the response text corresponding to the aforementioned prompt information.

[0162] The goal of using text prediction loss to correct the model parameters of the language model can be to minimize the text prediction loss (e.g., to 0), so that the subsequent language model can have a higher prediction confidence for the label text unit at the corresponding text position when generating each sample text unit at each sample time step.

[0163] The aforementioned sample prompts can be numerous. Following the same principle, this application can use a large number of sample prompts to perform multiple rounds of iterative training on the language model. After training is complete, the trained language model can be obtained. "Training complete" can mean that the model parameters of the language model have been trained to a convergent state, or that the number of training rounds of the language model has reached a set threshold. The number of training rounds refers to the number of iterative training rounds using sample prompts on the language model. When the model parameters of the language model have been trained to a convergent state, or when the number of training rounds has reached the set threshold, training is considered complete, and the trained language model is obtained.

[0164] Please refer to Figure 8, which is a schematic diagram of the principle of training a language model according to an embodiment of this application. As shown in Figure 8, the language model can include one shared network and three prediction networks (including prediction network 1 to prediction network 3). It can also have a total of three sample time steps, and the language model can associate N sample text units with each sample time step, including N sample text units associated with the first sample time step, N sample text units associated with the second sample time step, and N sample text units associated with the third sample time step.

[0165] By using N sample text units associated with the first sample time step, an intermediate prediction loss 1 can be generated for the language model at the first sample time step; an intermediate prediction loss 2 can be generated for the second sample time step using N sample text units associated with the second sample time step; and an intermediate prediction loss 3 can be generated for the third sample time step using N sample text units associated with the third sample time step. Therefore, by summing these three intermediate prediction losses, the overall text prediction loss of the language model can be obtained. This text prediction loss can be backpropagated to the language model to correct its parameters, ultimately resulting in the trained language model described above.

[0166] This application, through multi-token prediction, can not only predict the next token (such as the token at the next text position), but also predict multiple future tokens. This significantly improves the language model's ability to model complex contextual patterns, enhances its ability to capture long-distance dependencies and global semantic information, and allows the language model to focus on future contexts at longer distances (such as more distant text positions). Multi-token prediction can effectively capture global dependencies at the sentence, paragraph, and even document levels, thereby improving the training efficiency of the language model and the utilization rate of samples (such as sample prompts) by the language model, greatly enhancing the training effect of the language model.

[0167] This application achieves significant improvements over traditional language models in both the prediction objective and model structure (such as the introduction of N independent prediction networks) by constructing a multi-token prediction target, providing a reliable theoretical and technical foundation for efficient model training and high-quality text generation.

[0168] In actual experiments, this application also used four prediction heads (i.e., four prediction networks) to verify the proposed method. The text prediction speed was improved by 1.5 to 3 times compared to traditional language models, and the accuracy of text prediction was significantly improved. For traditional language models, one token is generated at each time step, and the text generation time increases linearly, with a text generation complexity of O(T), where T is the length of the generated text sequence. In contrast, this application generates multiple tokens at each time step, reducing the text generation complexity to O(T / N) and significantly reducing the text generation time. Text generation complexity is an indicator of the computational resources and time required by a language model in the process of generating text. Furthermore, in the code generation experiment (i.e., the text to be generated is code), compared to traditional language models, the language model of this application improved the accuracy of code generation by 7% on the HumanEval test set (a dataset of programming problems) and by 12% on the MBPP test set (a dataset of programming problems), thus demonstrating the effectiveness and practicality of the proposed method.

[0169] Please refer to Figure 9, which is a schematic diagram of the structure of a text prediction device provided in an embodiment of this application. As shown in Figure 9, the text prediction device 90 may include: an acquisition module 901, a prediction module 902, and a generation module 903.

[0170] The acquisition module 901 is used to acquire prompt information for text prediction.

[0171] The prediction module 902 is used to perform text prediction processing step by step according to the prompt information to obtain the target predicted text. The text prediction processing corresponding to each time step is used to generate N text units associated with each time step. Each text unit has its own position encoding information. If there are multiple time steps, there are K groups of text units with the same position encoding information among the N text units associated with two adjacent time steps. N is an integer greater than 1, K is a non-negative integer, and K is less than N.

[0172] The generation module 903 is used to generate response text corresponding to the prompt information based on the target predicted text; the response text contains text units generated by the text prediction processing corresponding to each time step.

[0173] In one implementation, if there is one time step, then the N text units associated with one time step are predicted based on a cue feature sequence, which is generated by performing feature transformation on the cue information; and,

[0174] If there are multiple time steps, there is a sequential order among the multiple time steps. The N text units associated with any time step are predicted based on the representation features and prompt feature sequence of the text units associated with the time steps before any time step. Among them, the representation features of any text unit are the features used to decode any text unit.

[0175] In one implementation, any one of at least one time step is the t-th time step, where t is a positive integer; the prediction module 902 performs text prediction processing step by step according to the prompt information to obtain the target predicted text, including:

[0176] Obtain the first predicted feature sequence; the first predicted feature sequence is composed of the representation features of text units in the first predicted text sequence, and the first predicted text sequence is composed of text units that have been predicted based on the prompt feature sequence at time steps before the t-th time step;

[0177] Based on the first predicted feature sequence and the prompt feature sequence, construct the first reference feature sequence corresponding to the t-th time step;

[0178] Text prediction is performed using the first reference feature sequence to generate N first text units associated at the t-th time step;

[0179] The second predicted text sequence is obtained based on N first text units, and the target predicted text is obtained through the second predicted text sequence.

[0180] In one implementation, the prediction module 902 obtains the target predicted text through a second predicted text sequence, including:

[0181] If the second predicted text sequence does not contain text units of the end type, then a second predicted feature sequence is constructed based on the representation features of the text units in the second predicted text sequence;

[0182] Based on the second predicted feature sequence and the cue feature sequence, construct the second reference feature sequence corresponding to the (t+1)th time step;

[0183] The second reference feature sequence is used for text prediction processing to generate N second text units associated at the (t+1)th time step;

[0184] The third predicted text sequence is obtained based on N second text units, and the target predicted text is obtained through the third predicted text sequence.

[0185] In one implementation, the prediction module 902 is further configured to:

[0186] If the second predicted text sequence contains text units of the end type, then the second predicted text sequence is used as the target predicted text;

[0187] Wherein, if there is at least one text unit following the end-of-text unit in the target predicted text, then the at least one text unit is a zero-value text unit supplemented after the end-of-text unit; the response text corresponding to the prompt information is generated based on the target predicted text, including:

[0188] The response text is obtained by removing end-of-text units and zero-value text units from the target prediction text.

[0189] In one embodiment, the position encoding information of any text unit is used to indicate the text position of any text unit in the reply text when any text unit is used in the reply text; the N text positions indicated by the N position encoding information of the N first text units are sequentially consecutive, and the N text positions indicated by the N position encoding information of the N second text units are sequentially consecutive.

[0190] Among them, there are K groups of text units with the same position encoding information in the N first text units and the N second text units; and the text position indicated by the position encoding information of the i-th text unit in the N first text units is located before the text position indicated by the position encoding information of the i-th text unit in the N second text units, where i is a positive integer and i is less than or equal to N.

[0191] In one implementation, the prediction module 902 obtains the second predicted text sequence based on N first text units, including:

[0192] Get the t×N text units associated with the t-th time step and the time steps before the t-th time step, where the t×N text units include N first text units;

[0193] From t×N text units, select M reference text units corresponding to positional encoding information respectively; the M positional encoding information includes the last positional encoding information and each positional encoding information before the text position indicated by the last positional encoding information. The last positional encoding information is the positional encoding information of the last position of the text position indicated by the N positional encoding information of the first text unit, and M is a positive integer.

[0194] The second predicted text sequence is constructed using M reference text units corresponding to M positional encoding information.

[0195] In one implementation, any one of the M location encoding information is the target location encoding information, and each of the t×N text units has its own prediction confidence; the prediction module 902 selects reference text units corresponding to the M location encoding information from the t×N text units in the following manner:

[0196] Extract one or more text units with target location encoding information from t×N text units;

[0197] The text unit with the highest prediction confidence among one or more text units is used as the reference text unit corresponding to the target location encoding information; or...

[0198] The text unit that appears most frequently among one or more text units is used as the reference text unit corresponding to the target location encoding information.

[0199] In one implementation, the prediction module 902 performs text prediction processing using a first reference feature sequence to generate N first text units associated with the t-th time step, including:

[0200] Obtain the trained language model; the trained language model contains a shared network and N prediction networks;

[0201] The shared network is invoked to learn features from the first reference feature sequence, generating shared sequence features of the first reference feature sequence;

[0202] Each prediction network is invoked to perform text prediction processing using shared sequence features, generating N first text units associated at the t-th time step;

[0203] One prediction network is used to generate a first text unit associated at time step t.

[0204] In one embodiment, the target predicted text is generated by calling a trained language model for text prediction processing; the text prediction device 90 further includes a training module 904, which is used for:

[0205] Obtain language model and sample prompt information;

[0206] The language model is invoked to perform text prediction processing based on sample prompts and sample time steps, generating N sample text units associated with each of the one or more sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text units associated with each sample time step.

[0207] Based on N sample text units associated with one or more sample time steps, generate the text prediction loss of the language model for the predicted sample text units;

[0208] The language model parameters are corrected using text prediction loss to obtain the trained language model.

[0209] In one implementation, N sample text units associated with one or more sample time steps are used to generate sample response text corresponding to sample prompt information. Each sample text unit associated with one or more sample time steps has its own position encoding information. The position encoding information of any sample text unit is used to indicate the text position of any sample text unit in the sample response text when any sample text unit is used in the sample response text.

[0210] The sample prompt information includes tag text, which serves as a reference for the sample response text. Each tag text contains L tag text units, and each of the L tag text units has position encoding information. The position encoding information of any tag text unit is used to indicate the text position of any tag text unit in the tag text, where L is a positive integer.

[0211] Any sample text unit associated with any sample time step has the same location encoding information as the label text unit to which it belongs among L label text units, and any sample text unit has a prediction confidence at any sample time step.

[0212] In one implementation, the training module 904 generates a method for the language model to calculate the text prediction loss for the predicted sample text units based on N sample text units associated with one or more sample time steps, including:

[0213] Based on the prediction confidence of the N sample text units associated with each sample time step, the intermediate prediction loss of the language model at each sample time step is generated.

[0214] The intermediate prediction losses of the language model at each sample time step are summed to generate the text prediction loss.

[0215] In one implementation, any one of one or more sample time steps is the s-th sample time step, where s is a positive integer; the training module 904 generates the intermediate prediction loss of the language model at each sample time step based on the prediction confidence of the N sample text units associated with each sample time step, including:

[0216] Based on the prediction confidence of each of the N sample text units associated with the s-th sample time step, generate the unit prediction loss corresponding to each sample text unit associated with the s-th sample time step.

[0217] The prediction losses of the N units corresponding to the N sample text units associated with the s-th sample time step are summed to generate the intermediate prediction loss of the language model at the s-th sample time step.

[0218] In one implementation, the prompt message is sent by the client; the generation module 903 is further configured to:

[0219] The reply text is returned to the client, allowing the client to output the reply text in its interface.

[0220] According to one embodiment of this application, the steps involved in the text prediction method shown in FIG3 can be executed by various modules in the text prediction device 90 shown in FIG9. For example, step S101 shown in FIG3 can be executed by the acquisition module 901 in FIG9, step S102 shown in FIG3 can be executed by the prediction module 902 in FIG9, and step S103 shown in FIG3 can be executed by the generation module 903 in FIG9.

[0221] This application can obtain prompt information for text prediction; and can perform text prediction processing step by step according to the prompt information to obtain the target predicted text; the text prediction processing corresponding to each time step is used to generate N text units associated with each time step, each text unit having its own positional encoding information. If there are multiple time steps, there are K groups of text units with the same positional encoding information among the N text units associated with two adjacent time steps, where N is an integer greater than 1, K is a non-negative integer, and K is less than N; and, it can also generate a response text corresponding to the prompt information according to the target predicted text; the response text contains the text units generated by the text prediction processing corresponding to each time step. Therefore, it can be seen that the method proposed in this application can generate multiple text units at each time step during the text prediction process. Furthermore, there can be K groups of text units with the same positional coding information among the N text units associated with each of two adjacent time steps. K can take any non-negative integer less than N. The larger the value of K, the more text units with the same positional coding information can be among the text units associated with each of two adjacent time steps, resulting in more correlation features between time steps and higher accuracy in text prediction. Conversely, the smaller the value of K, the fewer text units with the same positional coding information can be among the text units associated with each of two adjacent time steps, resulting in more text units with new positional coding information generated at each time step and higher efficiency in text prediction. It is evident that by generating multiple text units at each time step and setting a certain number (e.g., K groups) of text units with the same positional coding information between adjacent time steps, the efficiency and accuracy of text prediction can be improved using the method provided in this application.

[0222] According to one embodiment of this application, the various modules in the text prediction device 90 shown in FIG9 can be individually or entirely merged into one or more units, or some of the units can be further divided into multiple functionally smaller sub-units to achieve the same operation without affecting the technical effect of the embodiment of this application. The above modules are based on logical function division. In practical applications, the function of one module can also be implemented by multiple units, or the function of multiple modules can be implemented by one unit. In other embodiments of this application, the text prediction device 90 may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0223] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0224] According to one embodiment of this application, a computer program capable of executing the steps involved in the corresponding methods shown in the various embodiments of this application can be run on a general-purpose computer device (which may include processing elements and storage elements such as a central processing unit (CPU), random access storage medium (RAM), and read-only storage medium (ROM)) to construct the text prediction device 90 shown in FIG9. The aforementioned computer program can be recorded on a computer-readable recording medium and can be loaded into the aforementioned computer device via the computer-readable recording medium and run therein.

[0225] Please refer to Figure 10, which is a schematic diagram of the structure of a computer device provided in an embodiment of this application. As shown in Figure 10, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. Furthermore, in some embodiments, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. As shown in Figure 10, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.

[0226] In the computer device 1000 shown in Figure 10, the network interface 1004 provides network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0227] Obtain prompts for text prediction;

[0228] Based on the prompt information, text prediction processing is performed step by step according to time steps to obtain the target predicted text; the text prediction processing corresponding to each time step is used to generate N text units associated with each time step. Each text unit has its own position encoding information. If there are multiple time steps, there are K groups of text units with the same position encoding information among the N text units associated with each of two adjacent time steps. N is an integer greater than 1, K is a non-negative integer, and K is less than N.

[0229] The response text is generated based on the target predicted text, and contains text units generated by the text prediction processing at each time step.

[0230] It should be understood that the computer device 1000 described in the embodiments of this application can execute the text prediction method described in the various embodiments of this application, and can also execute the text prediction device 90 described in the embodiment corresponding to FIG9 above, which will not be repeated here. In addition, the beneficial effects of using the same method will not be repeated here.

[0231] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing a computer program. When a processor executes this computer program, it can perform the text prediction methods described in the various embodiments of this application; therefore, they will not be repeated here. Additionally, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer storage medium embodiments of this application, please refer to the description of the method embodiments of this application.

[0232] As an example, the aforementioned computer program can be deployed and executed on a single computer device, or deployed and executed on multiple computer devices located in one location, or executed on multiple computer devices distributed across multiple locations and interconnected via a communication network. These multiple computer devices distributed across multiple locations and interconnected via a communication network can form a blockchain network.

[0233] The aforementioned computer-readable storage medium can be an internal storage unit of the computer device, such as a hard drive or memory. It can also be an external storage device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, or flash card. Furthermore, the computer-readable storage medium can include both internal and external storage units of the computer device. This computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. It can also be used to temporarily store data that has been output or will be output.

[0234] This application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the text prediction methods described in the embodiments of this application; therefore, these descriptions will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments of this application, please refer to the description of the method embodiments of this application.

[0235] In summary, this application provides a text prediction method, apparatus, computer program product, computer device, and computer-readable storage medium. The computer device acquires prompt information for text prediction, and performs text prediction processing step by step according to the prompt information. The text prediction processing corresponding to each time step generates N text units associated with each time step. Each text unit has its own positional encoding information. If there are multiple time steps, there are K groups of text units with the same positional encoding information among the N text units associated with two adjacent time steps, where N is an integer greater than 1, K is a non-negative integer, and K is less than N. Finally, the target predicted text is obtained. Then, a response text corresponding to the prompt information is generated based on the target predicted text. The response text contains the text units generated by the text prediction processing corresponding to each time step. By generating multiple text units at each time step, the traditional method of generating single words one by one is changed, making the computational complexity of text generation change from the traditional linear growth to a lower complexity related to N, reducing the required computational resources and time, and significantly improving the efficiency of text prediction processing. Meanwhile, the existence of text units with the same positional encoding information between adjacent time steps enables the language model to learn the correlation features between the positions of preceding and following texts, enhancing its ability to understand and capture context, effectively reducing the cumulative error in the text generation process, and improving the accuracy of text prediction.

[0236] Furthermore, if there is one time step, the N text units associated with that time step are predicted based on the cue feature sequence, which is generated by transforming the cue information. This feature transformation converts the original information into a more regular and characteristic vector representation, facilitating efficient feature learning and analysis by the language model. This transformation process reduces information redundancy and noise, enabling the language model to more accurately grasp the key features of the cue information, providing a higher-quality data foundation for subsequent text prediction, thereby improving the accuracy and efficiency of prediction.

[0237] Furthermore, if there are multiple time steps with a sequential order, the N text units associated with any given time step are predicted based on the representation features and cue feature sequences of text units associated with previous time steps. The representation features of any given text unit are used to decode it. This prediction method, based on historical prediction results and cue feature sequences, allows the language model to fully utilize historical and current cue information to construct a more comprehensive contextual representation. By continuously integrating historical and current information, the language model reduces its reliance on single historical generation results, lowers prediction bias caused by the accumulation of historical errors, and thus improves the accuracy and consistency of text prediction.

[0238] Further, any one of at least one time step is designated as the t-th time step, where t is a positive integer. First, a first predicted feature sequence is obtained. This first predicted feature sequence consists of the representation features of text units in the first predicted text sequence, which in turn consists of text units predicted based on the cue feature sequence in time steps prior to the t-th time step. Next, based on the first predicted feature sequence and the cue feature sequence, a first reference feature sequence corresponding to the t-th time step is constructed. Then, text prediction processing is performed using the first reference feature sequence to generate N first text units associated with the t-th time step. Finally, a second predicted text sequence is obtained based on the N first text units, and the target predicted text is obtained through this second predicted text sequence. The process of constructing the first reference feature sequence effectively integrates historical prediction information and cue information, providing richer and more comprehensive information for text prediction at each time step. This information fusion enables the language model to more accurately capture the contextual semantics and structural features of the text, thereby generating text units more precisely and improving the accuracy and coherence of text prediction.

[0239] Furthermore, when obtaining the target predicted text through the second predicted text sequence, if the second predicted text sequence does not contain text units with end types, a second predicted feature sequence is constructed based on the representation features of the text units in the second predicted text sequence. Then, based on the second predicted feature sequence and the prompt feature sequence, a second reference feature sequence corresponding to the (t+1)th time step is constructed. Next, text prediction processing is performed using the second reference feature sequence to generate N second text units associated with the (t+1)th time step. Finally, a third predicted text sequence is obtained based on the N second text units, and the target predicted text is obtained through the third predicted text sequence. This method of dynamically adjusting subsequent prediction processes based on prediction results makes the text prediction process more adaptable and flexible. During the prediction process, the reference feature sequence is continuously updated, allowing the language model to make subsequent predictions based on the latest prediction information, ensuring that the prediction process can be continuously optimized until a predicted text sequence containing end-type text units is obtained, thereby improving the completeness and accuracy of the prediction.

[0240] Furthermore, if the second predicted text sequence contains text units of the ending type, then the second predicted text sequence is used as the target predicted text. If at least one text unit follows the ending type text unit in the target predicted text, then the at least one text unit is a zero-value text unit added after the ending type text unit. The ending type text units and zero-value text units in the target predicted text are removed to obtain the response text. Removing ending type text units and zero-value text units purifies the response text, removing useless information and ensuring that the response text contains only text units with practical meaning. This not only improves the conciseness and readability of the response text but also reduces the computational burden of subsequent processing and use of the response text, thus improving the efficiency of information transmission.

[0241] Furthermore, the positional encoding information of any text unit is used to indicate the text position of that text unit in the response text when it is used. The N text positions indicated by the N positional encoding information of the N first text units are sequentially consecutive, as are the N text positions indicated by the N positional encoding information of the N second text units. There are K groups of text units with the same positional encoding information among the N first text units and the N second text units, and the text position indicated by the positional encoding information of the i-th text unit in the N first text units precedes the text position indicated by the positional encoding information of the i-th text unit in the N second text units, where i is a positive integer and i is less than or equal to N. The setting of positional encoding information and the arrangement of text units provide the language model with explicit text positional information, enabling the language model to organize and generate text more accurately. This ordered text generation method ensures the structural and semantic coherence of the response text, reduces confusion and errors in text generation, and improves the quality and comprehensibility of the text.

[0242] Furthermore, when obtaining the second predicted text sequence based on N first text units, firstly, t×N text units associated with the t-th time step and the time steps preceding the t-th time step are obtained, where t×N text units include N first text units; then, from the t×N text units, M reference text units corresponding to positional encoding information are selected, where M positional encoding information includes the last positional encoding information and the positional encoding information for each text position preceding the text position indicated by the last positional encoding information. The last positional encoding information is the positional encoding information of the last text position among the N positional encoding information of the N first text units, where M is a positive integer; finally, the M reference text units corresponding to the M positional encoding information are used to construct the second predicted text sequence. The process of selecting reference text units from multiple text units is a filtering and optimization process. By selecting the most representative text units based on positional encoding information, redundant and inaccurate information can be removed, making the second predicted text sequence more refined and accurate. This helps the language model utilize information more efficiently in subsequent prediction processes, improving the accuracy and stability of predictions.

[0243] Furthermore, any one of the M positional encoding information is the target positional encoding information, and each text unit in the t×N text units has its own prediction confidence. When selecting reference text units corresponding to the M positional encoding information from the t×N text units, one or more text units with target positional encoding information are obtained from the t×N text units. The text unit with the highest prediction confidence among these one or more text units is selected as the reference text unit corresponding to the target positional encoding information, or the text unit with the highest frequency of occurrence among these one or more text units is selected as the reference text unit corresponding to the target positional encoding information. Selecting reference text units based on prediction confidence or frequency of occurrence ensures that the selected text units are the most reliable. Prediction confidence reflects the accuracy of the language model's prediction of the text unit, while frequency of occurrence reflects the stability and representativeness of the text unit. By selecting reference text units in this way, the introduction of erroneous and inaccurate text units is reduced, improving the accuracy and stability of the predicted text sequence.

[0244] Furthermore, when using the first reference feature sequence for text prediction processing to generate N first text units associated at time step t, a trained language model is first obtained. This trained language model includes a shared network and N prediction networks. Then, the shared network is invoked to learn features from the first reference feature sequence, generating shared sequence features. Next, each prediction network is invoked to perform text prediction processing using these shared sequence features, generating N first text units associated at time step t. One prediction network is used to generate one first text unit associated at time step t. The collaborative working mode of the shared network and the N prediction networks achieves the sharing and parallel processing of feature information. The shared network learns features from the first reference feature sequence, extracting general feature information, and the N prediction networks simultaneously perform text prediction based on this shared feature information. This parallel processing method significantly improves the efficiency of text prediction and reduces the time required for prediction. Simultaneously, the independent operation of multiple prediction networks helps reduce over-reliance on historical generation results, lowers accumulated errors, and improves the accuracy of text prediction.

[0245] Furthermore, the target predicted text is generated by calling the trained language model for text prediction processing. First, the language model and sample prompts are acquired. Then, the language model performs text prediction processing based on the sample prompts, according to the sample time steps, generating N sample text units associated with each of the one or more sample time steps. The text prediction processing corresponding to each sample time step is used to generate the N sample text units associated with that time step. Next, based on the N sample text units associated with each of the one or more sample time steps, the language model generates a text prediction loss for the predicted sample text units. Finally, the text prediction loss is used to correct the model parameters of the language model, resulting in the trained language model. By training and correcting the parameters of the language model using sample prompts and the text prediction loss, the language model can continuously learn and optimize. During training, the language model learns the features of different text patterns and contexts by processing a large amount of sample prompts. The text prediction loss serves as a feedback signal, guiding the adjustment of model parameters, enabling the model to gradually adapt to various text generation tasks and improving the generalization ability and accuracy of the language model.

[0246] Furthermore, N sample text units associated with one or more sample time steps are used to generate sample response text corresponding to the sample prompt information. Each sample text unit associated with one or more sample time steps has its own positional encoding information. The positional encoding information of any sample text unit is used to indicate the text position of any sample text unit in the sample response text when that sample text unit is used. The sample prompt information has label text, which serves as a reference for the sample response text. The label text contains L label text units, each of which has positional encoding information. The positional encoding information of any label text unit is used to indicate the text position of that label text unit in the label text, where L is a positive integer. Any sample text unit associated with any sample time step has the same positional encoding information as the label text unit to which it belongs among the L label text units. Each sample text unit has a prediction confidence at any sample time step. During training, the sample text units are mapped to the label text units, and the prediction confidence is introduced, providing a more accurate supervision signal for the training of the language model. By comparing sample text units and labeled text units, and analyzing prediction confidence, the prediction accuracy of the language model can be evaluated more accurately, thereby allowing for targeted adjustment of model parameters to improve training effectiveness and model accuracy.

[0247] Furthermore, when generating the text prediction loss of the language model for the predicted sample text units based on N sample text units associated with one or more sample time steps, the intermediate prediction loss of the language model at each sample time step is first generated based on the prediction confidence of the N sample text units associated with each sample time step. Then, the intermediate prediction losses of the language model at each sample time step are summed to generate the text prediction loss. Calculating and summing the intermediate prediction losses for each sample time step allows for a comprehensive and detailed measurement of the prediction bias of the language model throughout the training process. The intermediate prediction loss for each sample time step reflects the prediction accuracy at that time step. The text prediction loss obtained by summing all intermediate prediction losses can comprehensively evaluate the model's performance throughout the training process. This approach provides an accurate basis for adjusting model parameters, enabling the model to learn and optimize more effectively, thus improving the training effect.

[0248] Further, any one of the one or more sample time steps is designated as the s-th sample time step, where s is a positive integer. When generating the intermediate prediction loss of the language model at each sample time step based on the prediction confidence of the N sample text units associated with each sample time step, the unit prediction loss corresponding to each sample text unit associated with the s-th sample time step is first generated based on the prediction confidence of the N sample text units associated with the s-th sample time step. Then, the N unit prediction losses corresponding to the N sample text units associated with the s-th sample time step are summed to generate the intermediate prediction loss of the language model at the s-th sample time step. Calculating the unit prediction loss for each sample text unit and summing them to obtain the intermediate prediction loss allows for a more granular analysis of the language model's prediction performance at each sample time step. The unit prediction loss for each sample text unit reflects the model's prediction accuracy for that text unit. By summing all unit prediction losses, the model's performance at that sample time step can be evaluated more accurately. This detailed analysis provides more precise information for model optimization, contributing to further improvements in model performance.

[0249] The specification states that this application proposes a multi-token prediction approach, simultaneously predicting multiple future tokens at various time steps. This multi-token prediction method allows the language model to learn and capture more global semantic information from the context at each time step by leveraging the multiple tokens associated with previous time steps. Compared to traditional models, it reduces over-reliance on historical generation results, lowers accumulated errors, and thus significantly improves the accuracy of text prediction to generate text units. Simultaneously, because multiple future tokens can be predicted synchronously and in parallel at each time step, it accelerates text prediction speed and significantly improves efficiency. In natural language generation tasks, multi-token prediction reduces semantic drift in text generation, improves the coherence and consistency of generated text, and enhances the generalization ability of the language model for text prediction.

[0250] Furthermore, during language model training, this application utilizes multi-token prediction to predict not only the next token but also multiple future tokens, significantly improving the language model's ability to model complex contextual patterns and enhancing its ability to capture long-distance dependencies and global semantic information. Multi-token prediction effectively captures global dependencies at the sentence, paragraph, and even document levels, thereby improving the training efficiency and sample utilization of the language model, greatly enhancing the training effect. Simultaneously, by constructing the multi-token prediction objective, significant improvements over traditional language models are achieved in both prediction objectives and model structure, providing a reliable theoretical and technical foundation for efficient model training and high-quality text generation.

[0251] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0252] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0253] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A text prediction method, performed by a computer device, the method comprising: Obtain prompts for text prediction; Based on the aforementioned prompts, text prediction processing is performed step-by-step according to time steps to obtain the target predicted text. The text prediction processing corresponding to each time step generates N text units associated with that time step. Each text unit has its own positional encoding information. If there are multiple time steps, there are K groups of text units with the same positional encoding information among the N text units associated with two adjacent time steps, where N is an integer greater than 1, K is a non-negative integer, and K is less than N. The response text corresponding to the prompt information is generated based on the target predicted text; the response text contains text units generated by the text prediction processing corresponding to each time step.

2. The method as described in claim 1, wherein if the time step is one, then the N text units associated with one time step are predicted based on a cue feature sequence, wherein the cue feature sequence is generated by performing feature transformation on the cue information.

3. The method as described in claim 1 or 2, wherein if there are multiple time steps, the multiple time steps have a sequential order, and the N text units associated with any given time step are predicted based on the representation features of the text units associated with time steps preceding any given time step and the prompt feature sequence; wherein, The representation features of any of the text units are the features used to decode the text units.

4. The method according to any one of claims 1 to 3, wherein at least one of the time steps is the t-th time step, where t is a positive integer; the step of performing text prediction processing stepwise according to the prompt information to obtain the target predicted text includes: Obtain a first predicted feature sequence; the first predicted feature sequence is composed of the representation features of text units in the first predicted text sequence, and the first predicted text sequence is composed of text units that have been predicted based on the prompt feature sequence at time steps before the t-th time step; Based on the first predicted feature sequence and the prompt feature sequence, construct the first reference feature sequence corresponding to the t-th time step; The first reference feature sequence is used for text prediction processing to generate N first text units associated with the t-th time step; A second predicted text sequence is obtained based on the N first text units, and the target predicted text is obtained through the second predicted text sequence.

5. The method of claim 4, wherein obtaining the target predicted text through the second predicted text sequence comprises: If the second predicted text sequence does not contain text units of the end type, then a second predicted feature sequence is constructed based on the representation features of the text units in the second predicted text sequence; Based on the second predicted feature sequence and the prompt feature sequence, construct the second reference feature sequence corresponding to the (t+1)th time step; The second reference feature sequence is used for text prediction processing to generate N second text units associated with the (t+1)th time step; A third predicted text sequence is obtained based on the N second text units, and the target predicted text is obtained through the third predicted text sequence.

6. The method of claim 5, further comprising: If the second predicted text sequence contains text units of the ending type, then the second predicted text sequence is taken as the target predicted text; Wherein, if there is at least one text unit after the text unit of the ending type in the target predicted text, then the at least one text unit is a zero-value text unit supplemented after the text unit of the ending type; generating the reply text corresponding to the prompt information based on the target predicted text includes: The text units of the ending type and the zero-value text units in the target predicted text are removed to obtain the response text.

7. The method as described in claim 5 or 6, wherein the position encoding information of any of the text units is used to indicate the text position of any of the text units in the reply text when any of the text units is used in the reply text; the N text positions indicated by the N position encoding information of the N first text units are sequentially consecutive, and the N text positions indicated by the N position encoding information of the N second text units are sequentially consecutive; in, There are K groups of text units with the same position encoding information among the N first text units and the N second text units; and the text position indicated by the position encoding information of the i-th text unit among the N first text units is located before the text position indicated by the position encoding information of the i-th text unit among the N second text units, where i is a positive integer and i is less than or equal to N.

8. The method according to any one of claims 4 to 7, wherein obtaining the second predicted text sequence based on the N first text units comprises: Obtain t×N text units associated with the t-th time step and the time steps preceding the t-th time step, wherein the t×N text units include the N first text units; From the t×N text units, select M reference text units corresponding to positional encoding information respectively; the M positional encoding information includes the last positional encoding information and each positional encoding information before the text position indicated by the last positional encoding information, wherein the last positional encoding information is the last positional encoding information of the text position indicated by the N positional encoding information of the N first text units, and M is a positive integer; The second predicted text sequence is constructed using the M reference text units corresponding to the M positional encoding information.

9. The method as described in claim 8, wherein any one of the M location encoding information is target location encoding information, and each of the t×N text units has its own prediction confidence; the step of selecting reference text units corresponding to the M location encoding information from the t×N text units includes: Obtain one or more text units with the target location encoding information from the t×N text units; The text unit with the highest predicted confidence among the one or more text units is used as the reference text unit corresponding to the target location encoding information; or... The text unit that appears most frequently among the one or more text units is taken as the reference text unit corresponding to the target location encoding information.

10. The method according to any one of claims 4 to 9, wherein the step of performing text prediction processing using the first reference feature sequence to generate the N first text units associated at the t-th time step comprises: Obtain the trained language model; the trained language model includes a shared network and N prediction networks; The shared network is invoked to perform feature learning on the first reference feature sequence, generating shared sequence features of the first reference feature sequence; Each of the prediction networks is invoked to perform text prediction processing using the shared sequence features, generating the N first text units associated with the t-th time step; One of the prediction networks is used to generate a first text unit associated with the t-th time step.

11. The method according to any one of claims 1 to 10, wherein the target predicted text is generated by calling a trained language model for text prediction processing; the method further comprises: Obtain language model and sample prompt information; The language model is invoked to perform text prediction processing based on the sample prompt information according to the sample time step, generating one or more N sample text units associated with each of the sample time steps; the text prediction processing corresponding to each sample time step is used to generate N sample text units associated with each of the sample time steps. Based on N sample text units associated with one or more of the sample time steps, the language model generates a text prediction loss for the predicted sample text units. The text prediction loss is used to correct the model parameters of the language model, resulting in the trained language model.

12. The method of claim 11, wherein N sample text units associated with one or more of the sample time steps are used to generate sample response text corresponding to the sample prompt information, and each sample text unit associated with one or more of the sample time steps has its own position encoding information, and the position encoding information of any sample text unit is used to indicate the text position of any sample text unit in the sample response text when any sample text unit is used in the sample response text; in, The sample prompt information has tag text, which serves as a reference to the sample response text. The tag text contains L tag text units, each of which has position encoding information. The position encoding information of any tag text unit is used to indicate the text position of any tag text unit in the tag text, where L is a positive integer. Any sample text unit associated with any of the sample time steps has the same position encoding information as the label text unit to which it belongs among the L label text units, and any sample text unit has a prediction confidence at any of the sample time steps.

13. The method of claim 12, wherein generating the text prediction loss of the language model for the predicted sample text units based on N sample text units associated with one or more of the sample time steps comprises: Based on the prediction confidence of the N sample text units associated with each of the sample time steps, the intermediate prediction loss of the language model at each of the sample time steps is generated. The intermediate prediction loss of the language model at each sample time step is summed to generate the text prediction loss.

14. The method of claim 13, wherein any one of the one or more sample time steps is the s-th sample time step, where s is a positive integer; generating the intermediate prediction loss of the language model at each sample time step based on the prediction confidence of the N sample text units associated with each of the sample time steps includes: Based on the prediction confidence of each of the N sample text units associated with the s-th sample time step, generate the unit prediction loss corresponding to each sample text unit associated with the s-th sample time step; The prediction losses of the N units corresponding to the N sample text units associated with the s-th sample time step are summed to generate the intermediate prediction loss of the language model at the s-th sample time step.

15. A text prediction device, the device comprising: The acquisition module is used to acquire prompt information for text prediction. The prediction module is used to perform text prediction processing step by step according to the prompt information to obtain the target predicted text. The text prediction processing corresponding to each time step is used to generate N text units associated with each time step. Each text unit has its own position encoding information. If there are multiple time steps, there are K groups of text units with the same position encoding information among the N text units associated with two adjacent time steps. N is an integer greater than 1, K is a non-negative integer, and K is less than N. and The generation module is used to generate a response text corresponding to the prompt information based on the target predicted text; the response text contains text units generated by the text prediction processing corresponding to each time step.

16. A computer program product comprising a computer program stored in a computer-readable storage medium, the computer program being adapted to be read and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-14.

17. A computer device comprising a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1-14.

18. A computer-readable storage medium storing a computer program adapted to be loaded by a processor and to perform the steps of the method of any one of claims 1-14.