Text generation sequence length prediction model and training method thereof
By grafting the target network structure behind a large language model, the prediction model can accurately predict the sequence length when text generation, solving the problem that existing models cannot predict the number of future tokens, and improving the accuracy of the generation process.
Patent Information
- Application Number
- CN202510069697.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-06
AI Technical Summary
Existing large language models cannot effectively predict the number of tokens that need to be generated in the future when text generation.
A text-generated sequence length prediction model is designed, and the target network structure is grafted after the open source large language model, including a trainable request network, a Transformer network structure and a Linear network structure, is used to predict sequence length during the generation of tokens.
It realizes accurate prediction of sequence length during text generation, improving the prediction accuracy and practicality of the model.
Smart Images

Figure CN120106041A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a text generation sequence length prediction model and a training method thereof. Background Art
[0002] At present, large language models (LLMs) have made significant progress and have shown strong capabilities in the field of text generation. They usually have rich structures and parameters and can learn complex rules in text generation. However, when generating text with LLMs, it is impossible to predict how many tokens will be generated in the future. Summary of the invention
[0003] In order to overcome the shortcomings of the prior art, the present invention provides a text generation sequence length prediction model and a training method thereof, which can predict how many tokens will be generated in the future when LLM text is generated.
[0004] The first aspect of the present application provides a text generation sequence length prediction model, the text generation sequence length prediction model model includes: A preset open source large language model and a target network structure, wherein the target network structure is grafted onto the back of the open source large language model and is used to predict the number of tokens that need to be generated to complete the current conversation when the open source large language model generates tokens. The target network structure includes a trainable request network, a Transformer network structure, and a Linear network structure.
[0005] In an optional embodiment, the trainable request network structure regards the query vector in the multi-head attention layer as a trainable parameter. In an optional embodiment, the first sequence length of the output layer of the trainable request network structure is adjusted according to the maximum sequence length setting of the open source large language model. When the first sequence length is adjusted, the text generation sequence length prediction model is trained by acquiring a new training data set to adapt to the adjusted second sequence length, and the second sequence length is greater than the first sequence length.
[0006] In an optional embodiment, the Transformer network structure adopts RoPE position encoding processing by adding position encoding to the key vector and value vector.
[0007] In an optional implementation, the Linear network structure only retains the output linear layer, and converts the input data into output by performing a linear transformation through the output linear layer.
[0008] The second aspect of the present application provides a training method for a text generation sequence length prediction model, wherein the text generation sequence length prediction model includes a preset open source large language model and a target network structure, wherein the target network structure is grafted onto the back of the open source large language model, and the training method includes: When a text data set containing different conversation data is obtained, the same conversation data in the text data set is split to obtain multiple token data; Shuffling the multiple token data separated from the different conversation data to obtain a new data set; The target network structure is trained based on the new data set.
[0009] In an optional implementation, the target network structure includes a trainable request network structure, a Transformer network structure, and a Linear network structure, and the training method further includes: Determining a first prediction accuracy of the trainable request network structure, a second prediction accuracy of the Transformer network structure, and a third prediction accuracy of the Linear network structure; The first prediction accuracy, the second prediction accuracy, and the third prediction accuracy are compared to determine a network structure corresponding to the highest prediction accuracy in the target network structure.
[0010] In an optional embodiment, the training method further includes: Obtain an open source instruction dataset, where the open source instruction dataset is downloaded from an open source community and includes a user question part and a machine answer part; Regenerate the machine answer part in the open source instruction data set based on the open source large language model, and the correct value of the future sequence length corresponding to each token standard in the open source instruction data set to obtain a first labeled data set; The target network structure is trained based on the first labeled data set.
[0011] In an optional embodiment, the training method further includes: Acquire an instruction fine-tuning dataset, and use the instruction fine-tuning dataset to adjust the open source large language model; Annotating each token in the instruction fine-tuning data set with a correct value of the corresponding future sequence length to obtain a second annotated data set; The target network structure is trained based on the second labeled data set.
[0012] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the training method of the text generation sequence length prediction model when executing the computer program.
[0013] In summary, the text generation sequence length prediction model and its training method provided by the present application introduce a target network structure in the process of generating tokens by the open source large language model, which is grafted behind the open source large language model to predict the number of tokens that need to be generated to complete the current conversation while the open source large language model generates tokens. Among them, the target network structure can include a trainable request network structure, a Transformer network structure, and a Linear network structure. By combining the open source large language model and the target network structure, the sequence length can be predicted while the text is generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic diagram of the overall structure of a text generation sequence length prediction model shown in an embodiment of the present application; Figure 2 is a structural diagram of a QaT network structure shown in an embodiment of the present application; Figure 3 is a structural diagram of a Transformer network structure shown in an embodiment of the present application; Figure 4 It is a structural diagram of a Linear network structure shown in an embodiment of the present application; Figure 5 It is a flowchart of a training method for a text generation sequence length prediction model shown in an embodiment of the present application; Figure 6 1 is a schematic diagram of an evaluation of an EST header training result shown in an embodiment of the present application, wherein (a) is a distribution diagram of EST prediction errors, (b) is a distribution diagram of EST prediction error ratios, (c) is a distribution diagram of EST actual values, and (d) is a distribution diagram of EST prediction values; Figure 7 It is a structural schematic diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0015] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0016] The following will clearly and completely describe the concept, specific structure and technical effects of the present invention in combination with the embodiments and drawings, so as to fully understand the purpose, characteristics and effects of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, other embodiments obtained by technicians in this field without creative work, such as replacing a single-layer Transformer with a stack of multiple layers of Transformers, etc., all belong to the scope of protection of the present invention. In addition, all the connection / connection relationships involved in the patent do not refer to the direct connection of components, but refer to the fact that a better connection structure can be formed by adding or reducing connection accessories according to the specific implementation situation. The various technical features in the invention can be combined interactively without conflicting with each other.
[0017] Reference Figure 1 , which is a schematic diagram of the overall structure of a text generation sequence length prediction model shown in an embodiment of the present application.
[0018] Among them, the text generation sequence length prediction model includes a preset open source large language model and a target network structure grafted onto the back of the open source large language model. Specifically, the open source large language model is inserted into the position of the language model generation header, and the target network structure is inserted into the position of the EST header. The target network structure is an EST header network structure, including a trainable request (Query as Trainable parameters, QaT) network structure, a Transformer network structure, and a Linear network structure. The open source large language model is used to sample the ID of the next token according to the probability distribution. When the next token id is obtained, any target network structure is used to sample the expected sequence length according to the probability distribution to obtain the number of tokens to be generated.
[0019] In the embodiment of the present application, the preset open source large language model can be a conversational LLM in the prior art, wherein LLM is divided into a base class (e.g., a base model such as the Llama-3.2-3B model) and a conversation class (e.g., the Llama-3.2-3B-Instruct model).
[0020] In an optional embodiment, when the target network structure is a QaT network structure, the text generation sequence length prediction model is a dialogue-type LLM and a QaT network structure grafted onto the back of the dialogue-type LLM. The QaT network structure is a single-layer Transformer structure, and the query vector Q in the multi-head attention layer is regarded as a trainable parameter. The query vector Q can be regarded as a matrix such as (num_heads, head_dim), allowing the QaT network structure to directly learn the importance or preference of different attention heads, rather than relying on the dynamic transformation of the input data. Figure 2 , the input layer of the QaT network structure receives the last_hidden_state from the conversation class LLM, including the conversation data processed by the conversation class LLM. When the input layer receives the conversation data processed by the conversation class LLM, the linear layer performs a linear transformation on the input conversation data, and the normalization layer performs preprocessing on the conversation data to ensure that the input conversation data has a uniform scale and distribution. Specifically, the minimum-maximum normalization, Z-score normalization, etc. can be used to map the size of each numerical value in the data to a benchmark data distribution, thereby avoiding excessive, too small or invalid values in training. Furthermore, through the product layer provided by the feedforward neural network layer, the query vector Q, the key vector K and the value vector V calculate the attention weight through the dot product operation. Among them, softmax(QK T )V represents a typical attention operation. First, the dot product of the query vector Q and the key vector K is calculated, and then the softmax function is applied to convert the result into a probability distribution. Finally, the vector V is weighted and summed by the probability distribution to obtain the final attention output. And the activation function layer provided by the feedforward neural network layer introduces nonlinear transformations (for example, ReLU, sigmoid, tanh, etc.) to enable the neural network to learn complex functional relationships. Finally, the final prediction result is generated by the output layer, that is, the number of tokens that need to be generated to complete the current conversation data. In an embodiment of the present application, the output layer is a linear layer, which receives the feature representation from the previous layer and generates logits of EST values through linear transformation, and then uses the softmax function to convert the logits into a probability distribution of EST values.
[0021] In some embodiments, the output layer has 2048 neurons, which represent logits with integer values in the interval [0, 2048). The probability distribution of the predicted result value can be directly obtained through the softmax layer. The output result value represents how many tokens the LLM needs to generate to complete the current conversation. It should be noted that the output layer has 2048 neurons for illustration only, and is not a fixed value. It depends on the maximum sequence length setting of the LLM. If the sequence length is longer, a new data set is required for retraining.
[0022] The structural parameters of the QaT network structure may include, but are not limited to: the number of attention heads, the size of attention heads, the size of hidden layers, the size of hidden layers of feedforward neural networks, and the number of bins, etc. The values of the structural parameters are specifically referred to the following table:
[0023] In an optional embodiment, when the target network structure is a Transformer network structure, the text generation sequence length prediction model is a dialogue-class LLM and a Transformer network structure grafted onto the back of the dialogue-class LLM. Among them, the Transformer network structure can be a traditional Transformer structure, and its structural parameters and structural parameter values are consistent with those of the QaT network structure. In order to better compare with the QaT network structure to evaluate the impact of different position encoding processing methods on model performance, different from the LLM in the prior art that adds position encoding to the query vector Q and the key vector K, in the embodiment of the present application, the Transformer network structure adopts the RoPE position encoding processing method, by adding the position encoding to the key vector K and the value vector V. Specifically, in the multi-head attention mechanism, for each head, each original key vector K and value vector V is split into complex pairs according to dimension d (assuming d is an even number); the RoPE formula is applied to each complex pair to obtain a new complex pair; the rotated complex pairs are recombined into new key vectors K' and value vectors V'. Furthermore, the attention score is calculated using the modified key vector K' and the original query vector Q. Since the value vector V has also been processed by RoPE, the modified value vector V' will be used when calculating the weighted sum. The modified multi-head attention mechanism is integrated into each encoder layer and decoder layer in the Transformer network, so that each layer can use RoPE position encoding to process position information.
[0024] Refer to Figure 3, the input layer of the Transformer network structure also receives the last_hidden_state of the conversation class LLM, including the conversation data processed by the conversation class LLM. When the input layer receives the conversation data processed by the conversation class LLM, the linear layer performs a linear transformation on the input conversation data, that is, converts the last_hidden_state into a form suitable for subsequent processing. Specifically, the input data can be linearly combined through the weight matrix and the bias term to generate a new feature representation. And the conversation data can be preprocessed through the normalization layer to ensure that the input conversation data has a uniform scale and distribution. Specifically, the input conversation data can be converted into a standard form using minimum-maximum normalization, Z-score normalization, etc. Further, through the product layer provided by the feedforward neural network layer, the query vector Q, key vector K and value vector V calculate the attention weight through the dot product operation. Among them, softmax(QK T )V represents a typical attention operation. First, the dot product of the query vector Q and the key vector K is calculated, and then the softmax function is applied to convert the result into a probability distribution. Finally, the vector V is weighted and summed by the probability distribution to obtain the final attention output. And the activation function layer provided by the feedforward neural network layer introduces nonlinear transformations (for example, ReLU, sigmoid, tanh, etc.) to enable the neural network to learn complex functional relationships. Finally, the final prediction result is generated by the output layer, that is, the number of tokens that need to be generated to complete the current conversation data. In an embodiment of the present application, the output layer is a linear layer, which receives the feature representation from the previous layer and generates logits of EST values through linear transformation, and then uses the softmax function to convert the logits into a probability distribution of EST values.
[0025] In an optional implementation, when the target network structure is a linear network structure, the text generation sequence length prediction model is a dialogue-type LLM and a linear network structure grafted onto the back of the dialogue-type LLM, wherein the linear network structure is a linear network structure in a deep learning model. Figure 4 , the Linear network structure only retains the output linear layer, and the last_hidden_state received by the dialogue class LLM can be converted into output through linear transformation. The operation of the Linear network structure can be expressed as: y=wx+b, where w is the weight and b is the bias.
[0026] In some embodiments, the electronic device can be developed based on the AutoModelForCausalLMWithValueHead class of the trl library, replace the v_head of the class with the target network structure, and connect it after the LLM to achieve accurate prediction of the length of the text generation sequence.
[0027] Reference Figure 5 As shown, a flow chart of a method for training a text generation sequence length prediction model shown in an embodiment of the present application is provided, and the method for training a text generation sequence length prediction model includes the following steps.
[0028] S51, when a text data set containing different conversation data is obtained, the same conversation data in the text data set is split to obtain a plurality of token data.
[0029] After the target network structure is successfully grafted onto the dialogue class LLM to build the text generation sequence length prediction model, the text generation sequence length prediction model needs to be trained. Among them, the target network structure may include a QaT network structure, a Transformer network structure, and a Linear network structure. In some embodiments, the electronic device may collect a text data set containing different dialogue data, and split the same dialogue data in the text data set to obtain multiple token data. Among them, a token refers to an element in the text that can be regarded as a single unit, such as a word, number, or symbol. Specifically, the electronic device can cut the same dialogue data into smaller units such as words, words, or characters.
[0030] S52, shuffling the multiple token data separated from the different conversation data to obtain a new data set.
[0031] In some embodiments, the electronic device may shuffle the multiple token data after being split from different conversation data to disrupt their original order, increase the diversity of the data, and prevent the model from overfitting to a specific conversation mode or structure. The shuffled data will form a new data set (referred to as a new data set) for subsequent model training.
[0032] S53: Training the target network structure based on the new data set.
[0033] In some embodiments, before the new data set is input into the text generation sequence length prediction model, the new data set can be preprocessed, data normalized, encoded, etc. For the target network structure, it is necessary to ensure that its input format matches the output format of the open source large language model. Use the new data set to train the target network structure. During the training process, the open source large language model will be used as a feature extractor, and its output will be used as the input of the target network structure, and the mapping relationship from the output of the open source large language model to the length of the text generation sequence is learned through the target network structure. At the same time, a method is selected to measure the difference between the model predicted sequence length and the actual sequence length. In an embodiment of the present application, a cross entropy function is selected as the loss function, and an optimization algorithm is selected to minimize the loss function, such as stochastic gradient descent (SGD), Adam, etc. In other embodiments, the electronic device can use the TeDS (text level data segmentation) training strategy, that is, the correct results of the same conversation data are annotated token by token, and then all tokens of the same conversation data are treated as a group of data for processing using the high-speed parallel processing capability of the computing platform (Compute Unified Device Architecture, CUDA) launched by the graphics card manufacturer NVIDIA.
[0034] After the training is completed, the text generation sequence length prediction model can be deployed to actual application scenarios. When new conversation data is received, the trained text generation sequence length prediction model can be used for real-time prediction and output the predicted text generation sequence length.
[0035] In an optional embodiment, in order to verify the prediction accuracy of different target network structures for the text sequence length prediction model of the text generation sequence length, in the embodiment of the present application, all target network structures are verified on the gripper data set and the IW_1 data set, and the future token length of all test examples does not exceed 140. Among them, the prediction accuracy of the QaT network structure is called the first prediction accuracy, the prediction accuracy of the Transformer network structure is called the second prediction accuracy, and the prediction accuracy of the Linear network structure is called the third prediction accuracy. The performance indicators of different target network structures (QaT network structure, Transformer network structure and Linear network structure) on two data sets (gripper data set (also called gripper test set) and IW_1 data set (also called IW_1 test set)) are represented by the following table, including the average error values at different accuracy levels (Acc-10, Acc-50, Acc-100). By comparing them pairwise, the prediction accuracy of different target network structures for the text sequence length prediction model of the text generation sequence length can be determined. Among them, the gripper dataset is prepared for the robot self-programming task in the embodiment of the present application. Specifically, it is passed into Llama370B through some sample codes so that Llama3 70B can imitate the sample codes to give a series of similar data, which is divided into two parts: query and code. The IW_1 dataset is the v1 version of the Instruction-in-wild open source dataset.
[0036] Through comparative experiments, it is verified that the QaT network structure is the EST header structure with the highest prediction accuracy in the embodiments of the present application. The prediction accuracy of the traditional Transformer network structure is slightly lower than that of the QaT network structure, and the prediction accuracy of the Linear network structure is only about half of that of the QaT network structure. However, it is still applicable to scenarios where the prediction accuracy requirements are not high.
[0037]
[0038] In some embodiments, by randomly sampling 1020 data in the training set, the actual EST is compared with the model predicted EST, thereby obtaining a systematic evaluation of the prediction accuracy of the evaluation model. Figure 6 , where (a) is the distribution diagram of EST prediction error. The prediction error is obtained by subtracting the predicted value from the actual value. The overall distribution is beN (0.71, 9.52 2 ) is a normal distribution of EST prediction error ratio. Compared with the actual size of the error, the ratio of error length to actual length can better illustrate the EST prediction model, that is, the actual value of the target network structure. This ratio shows N (0.07, 0.542 ). (c) is the distribution of EST actual values. Most samples are distributed between [0, 80), which is consistent with the general application scenario of robot autonomous programming. (d) is the distribution of EST predicted values. It can be seen that the target network structure tends to misclassify the data in the three intervals of [20, 30), [50, 60), and [70, 80) into adjacent intervals.
[0039] according to Figure 6 It can be seen that the error of the target network structure for the sequence length presents a normal distribution centered on 0 as a whole, indicating that the target network structure can integrate the sequence semantic information to achieve a rough prediction of the future length of the sequence. Otherwise, the error distribution graph should present a certain random distribution.
[0040] In an optional embodiment, the training method further includes: Obtain an open source instruction dataset, where the open source instruction dataset is downloaded from an open source community and includes a user question part and a machine answer part; Regenerate the machine answer part in the open source instruction dataset based on the source large language model, and the correct value of the future sequence length corresponding to each token standard in the open source instruction dataset, to obtain a first labeled dataset; The target network structure is trained based on the first labeled data set.
[0041] In some embodiments, the electronic device can download an open source instruction data set from the open source community, where the open source instruction data set is presented in the form of "user question-machine answer", that is, it contains a large number of user questions and their corresponding machine answers. Then, the open source large language model pre-trained in the text generation sequence length prediction model is used as the target large model. In order to generate a machine answer corresponding to the open source instruction data set, the "user question" part in the open source instruction data set is input into the target large model, and generated when the temperature is 0, to ensure that the generated machine answer is more certain and consistent, and reduce randomness. Then, the electronic device can write a script to automatically mark the correct value of the future sequence length corresponding to each token to obtain the first labeled data set. Among them, the script will traverse each answer in the data set and calculate the length of the sequence after each token. The length value will be used as a label for subsequent model training. Finally, the electronic device can use the labeled data, that is, the first labeled data set to train the target network structure with ToDS (Token level Data Segmentation, a training strategy).
[0042] In an optional embodiment, the training method further includes: Acquire an instruction fine-tuning dataset, and use the instruction fine-tuning dataset to adjust the open source large language model; Annotating each token in the instruction fine-tuning data set with a correct value of the corresponding future sequence length to obtain a second annotated data set; The target network structure is trained based on the second labeled data set.
[0043] In some embodiments, it is assumed that there is a high-quality instruction fine-tuning dataset, which contains user instructions and corresponding machine answers for specific tasks or fields, and the instruction fine-tuning dataset is used to fine-tune the open source large language model in the text generation sequence length prediction model to improve the performance of the model on specific tasks. Similarly, the electronic device open source writes a script for automatically annotating the future sequence length corresponding to each token. Since the instruction fine-tuning dataset in this embodiment is already the standard answer, the electronic device can directly use the standard answer to calculate the sequence length after each token, and use it as a label to obtain a second annotated dataset. Next, use the annotated data, that is, the second annotated dataset to train the target network structure with ToDS.
[0044] The embodiment of the present application can realize the estimation of future computing resources, which can be used for LLM deployment parties to flexibly use limited computing resources to handle high-load conversation requests, and can also be used to display a certain progress bar for users in long text conversations. An example is listed below to describe the practical application of the embodiment of the present application. Currently 100 tokens have been generated. In the process of generating text, the electronic device can simultaneously use the target network structure (including QaT network structure, Transformer network structure and Linear network structure) in the text generation sequence length prediction model to predict the number of tokens that need to be generated to complete the current conversation. Assume that after 100 tokens are generated, the prediction model predicts that 45 tokens need to be generated to complete the conversation, that is, there are still 45 tokens to be generated.
[0045] See also Figure 7 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. In a preferred embodiment of the present application, the electronic device 7 includes a memory 71 , at least one processor 72 and at least one communication bus 73 .
[0046] Those skilled in the art should understand that Figure 7 The structure of the electronic device shown does not constitute a limitation of the embodiments of the present application, and can be either a bus structure or a star structure. The electronic device 7 can also include more or less other hardware or software than shown in the figure, or a different component arrangement.
[0047] In some embodiments, the electronic device 7 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices. The electronic device 7 may also include user equipment, which includes but is not limited to any electronic product that can interact with a user through a keyboard, mouse, remote control, touchpad, or voice control device, such as a personal computer, tablet computer, smart phone, digital camera, etc.
[0048] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0049] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0050] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A text generation sequence length prediction model, characterized in that: The text generation sequence length prediction model includes: A preset open source large language model and a target network structure, wherein the target network structure is grafted onto the back of the open source large language model and is used to predict the number of tokens that need to be generated to complete the current conversation when the open source large language model generates tokens. The target network structure includes a trainable request network structure, a Transformer network structure, and a Linear network structure.
2. The text generation sequence length prediction model according to claim 1, characterized in that: The trainable request network structure regards the query vector in the multi-head attention layer as a trainable parameter.
3. The text generation sequence length prediction model according to claim 1, characterized in that: The first sequence length of the output layer of the trainable request network structure is adjusted according to the maximum sequence length setting of the open source large language model. When the first sequence length is adjusted, the text generation sequence length prediction model is trained by acquiring a new training data set to adapt to the adjusted second sequence length, and the second sequence length is greater than the first sequence length.
4. The text generation sequence length prediction model according to claim 1, characterized in that: The Transformer network structure adopts the RoPE position encoding processing method by adding position encoding to the key vector and value vector.
5. The text generation sequence length prediction model according to claim 1, characterized in that: The linear network structure only retains the output linear layer, and transforms the input data into output by performing linear transformation through the output linear layer.
6. A training method for a text generation sequence length prediction model, characterized in that: The text generation sequence length prediction model includes a preset open source large language model and a target network structure, the target network structure is grafted onto the back of the open source large language model, and the training method includes: When a text data set containing different conversation data is obtained, the same conversation data in the text data set is split to obtain multiple token data; Shuffling the multiple token data separated from the different conversation data to obtain a new data set; The target network structure is trained based on the new data set.
7. The training method for a text generation sequence length prediction model according to claim 6, characterized in that: The target network structure includes a trainable request network structure, a Transformer network structure, and a Linear network structure, and the training method further includes: Determining a first prediction accuracy of the trainable request network structure, a second prediction accuracy of the Transformer network structure, and a third prediction accuracy of the Linear network structure; The first prediction accuracy, the second prediction accuracy, and the third prediction accuracy are compared to determine a network structure corresponding to the highest prediction accuracy in the target network structure.
8. The training method for a text generation sequence length prediction model according to any one of claims 6 to 7, characterized in that: The training method further comprises: Obtain an open source instruction dataset, where the open source instruction dataset is downloaded from an open source community and includes a user question part and a machine answer part; Regenerate the machine answer part in the open source instruction data set based on the open source large language model, and the correct value of the future sequence length corresponding to each token standard in the open source instruction data set to obtain a first labeled data set; The target network structure is trained based on the first labeled data set.
9. The training method for a text generation sequence length prediction model according to any one of claims 6 to 7, characterized in that: The training method further comprises: Acquire an instruction fine-tuning dataset, and use the instruction fine-tuning dataset to adjust the open source large language model; Annotating each token in the instruction fine-tuning data set with a correct value of the corresponding future sequence length to obtain a second annotated data set; The target network structure is trained based on the second labeled data set.
10. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the training method of the text generation sequence length prediction model described in any one of claims 6 to 9 are implemented.