Data prediction method, large prediction model training method, and related devices

By constructing multiple input data and using prediction models to perform target prediction tasks, the problem of slow inference speed of large language models is solved, and the effect of significantly improving data prediction efficiency is achieved.

CN119721215BActive Publication Date: 2025-06-27IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510242434.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-27
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

When responding to user requests, existing large language models only infer one data unit at a time, resulting in slow inference speed. How to improve the inference speed of large language models is a technical problem that needs to be solved urgently.

Method used

By constructing multiple input data, each input data includes currently predicted data and indicates the location of the data unit to be predicted, the prediction model is used to perform the target prediction task based on each input data, obtain multiple prediction data units, and add them to the currently predicted data until the data prediction completion conditions are met.

Benefits of technology

The efficiency of data prediction is significantly improved, so that the prediction model outputs multiple prediction data units at different locations at a time. Compared with the prior art, the method of outputting only one prediction data unit at a time improves the inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721215B_ABST
    Figure CN119721215B_ABST
Patent Text Reader

Abstract

The present application discloses a data prediction method, a prediction large model training method and related devices. The method includes: constructing n pieces of input data for this time, where each piece of input data includes currently predicted data and respectively indicates the position of a data unit to be predicted this time, and the positions indicated by each piece of input data are different, and n is greater than one; using the prediction large model to respectively execute the target prediction task based on each piece of input data to obtain n predicted data units; adding the n predicted data units to the currently predicted data to obtain the updated currently predicted data, where the position of each predicted data unit in the currently predicted data matches the position indicated by the corresponding input data sequence; repeating the above steps until the data prediction completion condition is satisfied, and obtaining the prediction result of the target prediction task based on the latest currently predicted data. By the above method, the present application can improve the prediction efficiency of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a data prediction method, a prediction large model training method, and related devices. Background Art

[0002] In recent years, large language models (LLMs) have demonstrated excellent performance in multiple application scenarios, such as natural language processing, machine translation, text generation, etc. Among them, large language models (large models) are used to reason about user requests to respond to user requests. In the prior art, when a large language model responds to a user request, it only infers one data unit each time, resulting in a slow inference speed of the large language model. Therefore, how to improve the inference speed of the large language model is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0003] The main technical problem to be solved by this application is to provide a data prediction method, a prediction large model training method, and related devices, which can improve the prediction efficiency of data.

[0004] To solve the above technical problem, a technical solution adopted by this application is: to provide a data prediction method, the method includes: constructing n pieces of input data for this time, where each piece of input data includes the currently predicted data and respectively indicates the position of a data unit to be predicted this time, the currently predicted data includes the data units predicted by the prediction large model when performing the target prediction task before this time, the positions indicated by each piece of input data are different, and n is greater than one; using the prediction large model to perform the target prediction task based on each piece of input data respectively to obtain n predicted data units; adding the n predicted data units to the currently predicted data to obtain the updated currently predicted data, where the positions of the predicted data units in the currently predicted data match the positions indicated by the corresponding input data sequences; repeating the above steps until the data prediction completion condition is met, and obtaining the prediction result of the target prediction task based on the latest currently predicted data.

[0005] Among them, in the n pieces of input data, the i-th piece of input data further includes i - 1 placeholders, where i is an integer in the closed interval of 1 to n; the positions indicated by each piece of input data are after the last element in the input data.

[0006] Among them, using the prediction large model to perform the target prediction task based on each piece of input data respectively to obtain n predicted data units includes: using the prediction large model to perform the target prediction task based on each piece of input data respectively to obtain n pieces of output data corresponding to the n pieces of input data respectively, where each piece of output data includes a predicted data unit corresponding to the position indicated by the input data.

[0007] Among the n pieces of output data, the j-th piece of output data includes the currently predicted data, j - 1 placeholders, and a predicted data unit, where j is an integer in the closed interval from 1 to n, and the predicted data unit is the last element of the output data.

[0008] Among them, the data prediction completion condition includes that the last data unit of the currently predicted data is an end symbol or the length of the currently predicted data reaches the upper limit value of the prediction length.

[0009] Among them, the currently predicted data also includes a prompt word, and the prompt word is used to prompt the prediction large model to perform the target prediction task; and / or, obtaining the prediction result of the target prediction task based on the latest currently predicted data, including: when the currently predicted data does not contain a prompt word, using the latest currently predicted data as the prediction result of the target prediction task; when the currently predicted data contains a prompt word, using the latest currently predicted data as the prediction result of the target prediction task, or, after deleting the prompt word in the latest currently predicted data, using it as the prediction result of the target prediction task.

[0010] Among them, the method further includes the following steps for training the prediction large model: obtaining original training data, where the original training data includes multiple sample data units; dividing the multiple sample data units into blocks to obtain several data blocks, where the several data blocks include at least one standard data block, and the standard data block includes n sample data units; selecting at least one target data block from the several data blocks, where the at least one target data block includes the standard data block; for each target data block, constructing m - 1 new training data of the target data block, where m is the number of sample data units in the target data block, and the new training data is obtained by replacing the target data block in the original training data with a replacement data block, and the replacement data blocks in the m - 1 new training data are respectively composed of 1 to m - 1 placeholders; training the prediction large model using the original training data and the new training data of each target data block.

[0011] Among them, the original training data and the new training data include sample prompt words, and the sample prompt words are used to prompt the prediction large model to perform sample prediction tasks based on the original training data or the new training data; and / or, all data blocks except the last data block among the several data blocks are standard data blocks, and the last data block is a standard data block or the number of sample data units is less than n; and / or, selecting at least one target data block from the several data blocks includes: taking all data blocks as target data blocks.

[0012] To solve the above technical problems, another technical solution adopted in this application is: to provide a prediction large model training method, which includes: obtaining original training data, where the original training data includes multiple sample data units; dividing the multiple sample data units into blocks to obtain several data blocks, where the several data blocks include at least one standard data block, and the standard data block includes n sample data units; selecting at least one target data block from the several data blocks, where the at least one target data block includes the standard data block; for each target data block, constructing m - 1 new training data of the target data block, where m is the number of sample data units in the target data block, and the new training data is obtained by replacing the target data block in the original training data with a replacement data block, and the replacement data blocks in the m - 1 new training data are respectively composed of 1 to m - 1 placeholders; using the original training data and the new training data of each target data block to train the prediction large model.

[0013] To solve the above technical problems, yet another technical solution adopted in this application is: to provide a data prediction device, including: a construction module, a task execution module, a data update module, and a prediction module. The construction module is used to construct n input data for this time, where each input data includes the currently predicted data and respectively indicates the position of a data unit to be predicted this time. The currently predicted data includes the data units predicted by the prediction large model when performing the target prediction task before this time, and the positions indicated by each input data are different, and n is greater than one; the task execution module is used to use the prediction large model to respectively execute the target prediction task based on each input data to obtain n predicted data units; the data update module is used to add the n predicted data units to the currently predicted data to obtain the updated currently predicted data, where the positions of the predicted data units in the currently predicted data match the positions indicated by the corresponding input data sequences; the prediction module is used to repeat the above steps until the data prediction completion condition is met, and obtain the prediction result of the target prediction task based on the latest currently predicted data.

[0014] To solve the above technical problems, another technical solution adopted in this application is: to provide a prediction large model training device, including: an acquisition module, a data chunking module, a data selection module, a data construction module, and a training module; the acquisition module is used to acquire original training data, where the original training data includes multiple sample data units; the data chunking module is used to chunk the multiple sample data units to obtain several data chunks, where the several data chunks include at least one standard data chunk, and the standard data chunk includes n sample data units; the data selection module is used to select at least one target data chunk from the several data chunks, where the at least one target data chunk includes the standard data chunk; the data construction module is used to construct m - 1 new training data for each target data chunk, where m is the number of sample data units in the target data chunk, and the new training data is obtained by replacing the target data chunk in the original training data with a replacement data chunk, and the replacement data chunks in the m - 1 new training data are respectively composed of 1 to m - 1 placeholders; the training module is used to train the prediction large model by using the original training data and the new training data of each target data chunk.

[0015] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, including a memory and a processor that are coupled to each other, and the memory stores program instructions; the processor is used to execute the program instructions stored in the memory to implement the above method.

[0016] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium for storing program instructions, and the program instructions can be executed to implement the above method.

[0017] In the above solution, the number of input data for the prediction large model each time is n, and each input data includes the position of a data unit to be predicted this time, and the positions indicated by each input data are different. Therefore, when the prediction large model performs the target prediction task based on each input data respectively, n predicted data units at different positions can be predicted based on each input data. Compared with the existing method where the prediction large model only outputs one predicted data unit each time, the method of constructing the n input data this time in this application to make the prediction large model output n predicted data units at different positions each time can significantly improve the efficiency of data prediction. Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of an embodiment of the data prediction method provided by this application;

[0019] Figure 2 It is a schematic flowchart of an embodiment of the prediction large model training method provided by this application;

[0020] Figure 3 It is a schematic framework diagram of an embodiment of the data prediction device provided by this application;

[0021] Figure 4 It is a schematic framework diagram of an embodiment of the prediction large model training device provided by this application;

[0022] Figure 5 It is a schematic framework diagram of an embodiment of the electronic device provided by this application;

[0023] Figure 6 It is a schematic framework diagram of the computer-readable storage medium provided by this application. Detailed implementation manners

[0024] To make the purpose, technical solutions and effects of this application clearer and more definite, the following further describes this application in detail with reference to the accompanying drawings and by way of examples.

[0025] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0026] It should be noted that the prediction large model generally performs reasoning in an autoregressive manner. The so-called autoregressive manner means that at any time step , given the input sequence P, , ,... When (where P represents the prompt word), the model needs to predict . In this way, when the prediction large model needs to predict a sequence of length n, a total of n + 1 time steps are required, and the last additional time step is used to predict the end of the sequence, that is, EOS.

[0027] Among them, in this method, when the prediction large model makes a prediction at each time step (each time), only one data unit is predicted, resulting in low prediction efficiency of the prediction large model. In order to improve the data prediction efficiency of the prediction large model, the present application creatively proposes to construct multiple input data for the prediction large model at each time step (each time), and each piece of input data includes the currently predicted data and the positions respectively indicating a data unit to be predicted this time, so that the prediction large model can predict the same number of data units as the number of input data at each time step, thereby improving the data prediction efficiency of the prediction large model.

[0028] Specifically, please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the data prediction method provided by the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment includes:

[0029] S11: Construct n pieces of input data for this time, where each piece of input data includes the currently predicted data and the positions respectively indicating a data unit to be predicted this time. The currently predicted data includes the data units predicted by the prediction large model when performing the target prediction task before this time. The positions indicated by each piece of input data are different, and n is greater than one.

[0030] It should be noted that when the prediction large model of this embodiment performs the target prediction task, it is generally carried out in an autoregressive manner. That is, after predicting n predicted data units based on the n pieces of input data constructed this time, the n predicted data units are added to the currently predicted data to obtain the updated currently predicted data, and steps S11 - S14 are repeatedly executed until the data prediction completion condition is met.

[0031] Specifically, the updated currently predicted data can be used to construct n pieces of input data for the next time, and the prediction large model is used to perform the target prediction task based on each newly constructed piece of input data respectively to obtain n predicted data units. The n predicted data units are added to the currently predicted data to obtain the updated currently predicted data until the data prediction completion condition is met, and the prediction result of the target prediction task is obtained based on the latest currently predicted data.

[0032] Among them, the data prediction completion condition includes that the last data unit of the currently predicted data is an end symbol or the length of the currently predicted data reaches the prediction length upper limit value. The end symbol can be any symbol indicating the end of the prediction, for example, it is <eos>, the specific upper limit value of the prediction length can be preset according to actual needs.

[0033] In some embodiments, among the n pieces of input data constructed, the i-th piece of input data further includes i - 1 placeholders, that is, the number of placeholders included in each piece of input data is different, where i is an integer in the closed interval of 1 to n; the position indicated by each piece of input data is after the last element in the input data. That is, in this embodiment, the position of a data unit to be predicted is after the last element in the input data.

[0034] For example, the n pieces of input data constructed for the first time are as follows:

[0035] ="Please help me write a poem:”

[0036] ="Please help me write a poem: <m>”

[0037] ="Please help me write a poem: <m> <m>”

[0038] ="Please help me write a poem: <m> <m> <m>”

[0039] ="Please help me write a poem: <m> <m> <m> <m>”

[0040] ="Please help me write a poem: <m> <m> <m> <m> <m>”

[0041] Among them, "Please help me write a poem" is a prompt, which is used to prompt the prediction large model to predict a poem; <m>Represents a placeholder, the first input data in the example ="Please help me write a poem:” without placeholder, the second input data ="Please help me write a poem: <m>” including a placeholder <m>, The third input data ="Please help me write a poem: <m> <m> <m>”Include 3 placeholders <m>, and the position indicated by each input data is after the last element in the corresponding input data, where each placeholder in each input data <m>is an element in the input data.

[0042] It should be noted that the placeholder is used to tell the prediction large model that there is a data unit at the position occupied by the placeholder, and to tell the prediction large model the position of the data unit to be predicted. In one embodiment, the position of the data unit predicted by the prediction large model is behind the last placeholder in the input data.

[0043] Of course, in other embodiments, among the n pieces of input data constructed, the i-th piece of input data may include i placeholders, where i is an integer in the closed interval from 1 to n; the position indicated by each piece of input data is the last element in the corresponding input data, that is, the position indicated by each piece of input data is the position where the last placeholder in the corresponding input data is located.

[0044] S12: Use the prediction large model to perform the target prediction task based on each piece of input data respectively, and obtain n predicted data units.

[0045] In this embodiment, the target prediction task may be, but is not limited to, natural language processing, machine translation, text generation, etc., and may also be image generation, etc. The specific data unit is related to the target prediction task. For example, if the target prediction task is related to text generation, the data unit may be a character in the text (such as a word, phrase, or symbol), and if the target prediction task is related to image generation, the data unit may be a pixel point in the image or a frame of the image. The specific data unit can be predefined according to actual needs.

[0046] For the convenience of description, the following explanations of the solutions of this application are all given by taking the target prediction task related to text generation as an example, and the protection scope of this application cannot be limited thereby.

[0047] In one implementation scenario, the prediction large model has the ability to handle multiple tasks. In order to make the model aware of the target prediction task to be executed, when constructing the n pieces of input data for the prediction large model for the first time or each time, the constructed input data may include prompt words to use the prompt words to prompt the target prediction task that the prediction large model needs to execute.

[0048] In one embodiment, the currently predicted data in each piece of input data includes prompt words for prompting the target prediction task that the prediction large model needs to execute.

[0049] In some embodiments, step S12 uses the prediction large model to perform the target prediction task based on each piece of input data respectively, and obtain n predicted data units, including: using the prediction large model to perform the target prediction task based on each piece of input data respectively, and obtain n pieces of output data corresponding to the n pieces of input data respectively, where each piece of output data includes a predicted data unit corresponding to the position indicated by the input data.

[0050] Generally speaking, in this embodiment, the number of constructed input data is the same as the number of output data of the prediction large model, and each output data includes a predicted data unit corresponding to the position indicated by the input data.

[0051] In some embodiments, among the n output data output by the prediction large model, the j-th output data includes the currently predicted data, j - 1 placeholders, and a predicted data unit, where j is an integer in the closed interval of 1 to n, and the predicted data unit is the last element of the output data.

[0052] Exemplarily, the n output data output by the prediction large model are as follows:

[0053] ="Please help me write a poem: White"

[0054] ="Please help me write a poem: <m>day

[0055] ="Please help me write a poem: <m> <m>According to "

[0056] ="Please help me write a poem: <m> <m> <m>Mountain

[0057] ="Please help me write a poem: <m> <m> <m> <m>exhaust; all; to the greatest extent

[0058] ="Please help me write a poem: <m> <m> <m> <m> <m>,”

[0059] Among them, please help me write a poem as a prompt. The currently predicted data includes the prompt. For the first output data ="Please help me write a poem: White", and the predicted data unit is "White". For the second output data ="Please help me write a poem: <m>"day", the prediction data unit is "day", placeholder <m>The quantity is 1, for the fifth output data ="Please help me write a poem: <m> <m> <m> <m>"exhaust", the prediction data unit is "exhaust", placeholder <m>The quantity is 4.

[0060] S13: Add n predicted data units to the currently predicted data to obtain the updated currently predicted data, where the positions of the respective predicted data units in the currently predicted data match the positions indicated by the corresponding input data sequences.

[0061] Exemplarily, the n predicted data units output by the prediction large model are respectively "white", "day", "lean", "on", "the", "hill", and ",", and add "The sun along the mountain bows" to the currently predicted data to obtain the updated currently predicted data "Please help me write a poem: The sun along the mountain bows".

[0062] ="Please help me write a poem: white"

[0063] ="Please help me write a poem: <m>day

[0064] ="Please help me write a poem: <m> <m>According to "

[0065] ="Please help me write a poem: <m> <m> <m>Mountain

[0066] ="Please help me write a poem: <m> <m> <m> <m>Exhaustion

[0067] ="Please help me write a poem: <m> <m> <m> <m> <m>,”

[0068] Further, using the currently predicted data "Please write a poem for me: The sun along the mountain bows," construct n new input data as follows:

[0069] ="Please write a poem for me: The sun along the mountain bows,"

[0070] ="Please write a poem for me: The sun along the mountain bows, <m>”

[0071] ="Please help me write a poem: As the sun sets behind the mountains, <m> <m>”

[0072] ="Please help me write a poem: The sun along the mountain bows, <m> <m> <m>”

[0073] ="Please help me write a poem: The sun along the mountain bows," <m> <m> <m> <m>”

[0074] ="Please help me write a poem: The sun along the mountain bows," <m> <m> <m> <m> <m>”

[0075] S14: Repeat the above steps until the data prediction completion condition is met, and obtain the prediction result of the target prediction task based on the latest currently predicted data.

[0076] In some embodiments, step S14 obtaining the prediction result of the target prediction task based on the latest currently predicted data includes the following two cases:

[0077] Case 1: When the currently predicted data does not contain a prompt word, use the latest currently predicted data as the prediction result of the target prediction task.

[0078] Case 2: When the currently predicted data contains a prompt word, use the latest currently predicted data as the prediction result of the target prediction task, or, after deleting the prompt word in the latest currently predicted data, use it as the prediction result of the target prediction task.

[0079] Generally speaking, in Case 2, it is hoped that the prediction result of the final target prediction task may or may not contain a prompt word.

[0080] Among them, before step S12 using the prediction large model to execute the target prediction task based on each of the input data and obtaining n prediction data units, the prediction large model also needs to be trained.

[0081] Specifically, please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the prediction large model training method provided by this application. As Figure 2 shown, the prediction large model training method includes:

[0082] S21: Obtain original training data, where the original training data includes multiple sample data units.

[0083] The sample data unit in this embodiment can be understood as a token, such as a character, a pixel point or a frame of image, and the specific sample data unit can be defined in advance.

[0084] For example, the original training data is: X = P, , ,... , where P represents the sample prompt word, , ,... represents a sequence with a length of n of the expected output, , ,... each of which is a sample data unit.

[0085] S22: Divide multiple sample data units into chunks to obtain a number of data chunks, where the number of data chunks includes at least one standard data chunk, and the standard data chunk includes n sample data units.

[0086] In this embodiment, the number of data chunks includes at least one standard data chunk, and the standard data chunk includes n sample data units.

[0087] In one embodiment, data chunks other than the last data chunk among the number of data chunks are all standard data chunks, and the last data chunk is a standard data chunk or the number of sample data units is less than n. Generally speaking, in this embodiment, the number of data chunks is obtained by evenly dividing multiple sample data units. For example, divide the number of sample data units in the expected output sequence by n to obtain each data chunk after division; among them, the number of sample data units in the last data chunk is less than or equal to n. If the number of sample data units in the last data chunk is n, then the last data chunk is also a standard data chunk.

[0088] For example, the original training data is P = "Please help me write a poem: ", ="Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright; Bowing, in homesickness I'm drowned" <eos>”. The expected output sequence includes a total of 24 sample data units. The number of sample data units included in the expected standard data block is 6. Therefore, each data block obtained by evenly dividing multiple sample data units is as follows: Data block 1 = "Before my bed a pool of light,", Data block 2 = "I wonder if it's frost aground.", Data block 3 = "Looking up, I find the moon bright;", Data block 4 = "Bowing, in homesickness I'm drowned." <eos>”.

[0089] In this example, each line of poetry, with a comma at the end of the line, or <eos>, which is exactly 6 characters, that is, the data blocks obtained for each sub-block are all standard data blocks.

[0090] In one embodiment, the number of sample data units expected to be included in the standard data block can be preset to a fixed value, or a range can be preset, and then a suitable value can be selected from this range using a related algorithm. For example, the best value can be searched from the set range using the method of hyperparameter search.

[0091] It should be noted that the method of evenly dividing multiple sample data units into several data blocks as described above can make the number of sample data units included in each of the remaining data blocks except the last data block the same, which is beneficial to the efficient utilization of computing resources.

[0092] Of course, in other embodiments, when the utilization efficiency of computing resources is within an acceptable range, the above-mentioned method of evenly dividing blocks may not be used to obtain each data block. For example, a random block division method can be used for block processing.

[0093] S23: Select at least one target data block from several data blocks, where at least one target data block includes a standard data block.

[0094] The purpose of selecting the target data block is to use the selected target data block to construct new training data to expand the quantity of training data. Therefore, in order to obtain more new training data subsequently, all data blocks can be used as target data blocks.

[0095] Since the prediction large model is generally fine-tuned based on the existing model, and the prediction large model has the ability to handle multiple tasks, in order to enable the prediction large model to be trained to clearly identify the task to be executed, sample prompt words can be set in the original training data and the newly added new training data to prompt the prediction large model to perform sample prediction tasks based on the original training data or new training data.

[0096] S24: For each target data block, construct m - 1 new training data of the target data block, where m is the number of sample data units in the target data block. The new training data is obtained by replacing the target data block in the original training data with a replacement data block, and the replacement data blocks in the m - 1 new training data are respectively composed of 1 to m - 1 placeholders.

[0097] For example, taking the above data block 1 = "Before my bed a pool of light," as the target data block, the m - 1 new training data of the target data block are constructed as follows:

[0098] = "Please help me write a poem: <m>Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright; Bowing, in homesickness I'm drowned. <eos>”

[0099] ="Please help me write a poem: <m> <m>The bright moonlight seems like frost on the ground. I raise my head to gaze at the bright moon, and lower my head, thinking of my hometown. <eos>”

[0100] ="Please help me write a poem: <m> <m> <m>The moonlight seems like frost on the ground. I raise my head to look at the bright moon, and lower my head to think of my hometown. <eos>”

[0101] ="Please help me write a poem: <m> <m> <m> <m>The moonlight seems like frost on the ground. I raise my head and look at the bright moon, then lower my head and think of my hometown. <eos>”

[0102] ="Please help me write a poem: <m> <m> <m> <m> <m>, I suspected it was frost on the ground. I looked up at the bright moon and lowered my head, thinking of my hometown. <eos>”

[0103] Among them, m represents the number of sample data units in the target data block "In front of my bed the moonlight glows", specifically 6. The number of new training data constructed based on this target data block is 5. Among them, each new training data is the original training data "P=" Please help me write a poem: =”In front of my bed the moonlight glows, I wonder if it's frost on the ground. I raise my head to look at the bright moon, Lower my head and think of home <eos>"The target data block in "Before my bed a pool of light," is replaced with a replacement data block. Specifically, it is replaced with " <m>Before the bright moon shines, "", "" <m> <m>Bright moonlight, "," <m> <m> <m>Moonlight, ”, " <m> <m> <m> <m>"Light" and " <m> <m> <m> <m> <m>,”.

[0104] Any one of the above data blocks 2-4 is used as the target data block, and the method of constructing the corresponding new training data is the same as that of data block 1. Taking data block 4 as an example, the newly constructed training data is as follows:

[0105] ="Please help me write a poem: Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright, <m>Missing my hometown <eos>”

[0106] ="Please help me write a poem: Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright, <m> <m>Missing My Hometown <eos>”

[0107] ="Please help me write a poem: Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright, <m> <m> <m>Hometown <eos>”

[0108] ="Please help me write a poem: Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright, <m> <m> <m> <m>township <eos>”

[0109] ="Please help me write a poem: Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright, <m> <m> <m> <m> <m> <eos>”

[0110] S25: Train the prediction large model using the original training data and the new training data of each target data block.

[0111] After constructing the new training data of each target data block as described above, for each original training data, use the original training data and the new training data constructed based on each target data block in the original training data as a training sample data to train the prediction large model. During the actual training process, multiple training sample data need to be constructed to train the prediction large model to convergence using multiple training sample data.

[0112] Among them, each training sample data contains several pieces of training data, and the corresponding prediction large model will also output the same number of prediction data as the training data (predicting what the data unit is after the last element in the input data). The difference between the prediction data and the sample annotation of each piece of training data in the training sample data can be synthesized to obtain a comprehensive difference, and then the network parameters of the prediction large model can be adjusted using this comprehensive difference.

[0113] It should be noted that for each data block in each original training data in the prediction large model training method provided in this application, multiple (m - 1) corresponding new training data will be constructed based on this data block, and the replacement data blocks in each new training data are respectively composed of 1 to m - 1 placeholders, that is, the number of placeholders in each new training data constructed based on each data block is different, and the positions of the sample data units that need to be predicted by the prediction large model based on each piece of training data are also different. It can be understood that for each target data block, using the new training data constructed based on this target data block to train the prediction large model can train the prediction large model's ability to predict data units at different positions in this target data block in parallel.

[0114] In the above solution, the number of input data for the prediction large model each time is n, and each piece of input data includes an indication of the position of a data unit to be predicted this time, and the positions indicated by each piece of input data are different. Therefore, when using the prediction large model to perform the target prediction task based on each piece of input data respectively, n prediction data units at different positions can be predicted based on each piece of input data. Compared with the existing method where the prediction large model only outputs one prediction data unit each time, the method of constructing the n pieces of input data this time in this application so that the prediction large model outputs n prediction data units at different positions each time can significantly improve the efficiency of data prediction.

[0115] Please refer to Figure 3 , Figure 3 It is a schematic framework diagram of an embodiment of the data prediction device provided by this application. In this embodiment, the data prediction device 30 includes a construction module 31, a task execution module 32, a data update module 33, and a prediction module 34. The construction module 31 is used to construct n pieces of input data for this time. Among them, each piece of input data includes the currently predicted data and respectively indicates the position of a data unit to be predicted this time. The currently predicted data includes the data units predicted by the prediction large model when performing the target prediction task before this time. The positions indicated by each piece of input data are different, and n is greater than one; the task execution module 32 is used to use the prediction large model to perform the target prediction task based on each piece of input data respectively, and obtain n predicted data units; the data update module 33 is used to add the n predicted data units to the currently predicted data to obtain the updated currently predicted data, where the positions of the predicted data units in the currently predicted data match the positions indicated by the corresponding input data sequences; the prediction module 34 is used to repeat the above steps until the data prediction completion condition is met, and obtain the prediction result of the target prediction task based on the latest currently predicted data.

[0116] In some embodiments, among the n pieces of input data constructed by the construction module 31, the i-th piece of input data further includes i - 1 placeholders, where i is an integer within the closed interval of 1 to n; the positions indicated by each piece of input data are after the last element in the input data.

[0117] In some embodiments, the task execution module 32 uses the prediction large model to perform the target prediction task based on each piece of input data respectively, and obtain n predicted data units, including: using the prediction large model to perform the target prediction task based on each piece of input data respectively, and obtain n pieces of output data corresponding to the n pieces of input data respectively, where each piece of output data includes a predicted data unit corresponding to the position indicated by the input data.

[0118] In some embodiments, among the n pieces of output data, the j-th piece of output data includes the currently predicted data, j - 1 placeholders, and a predicted data unit, where j is an integer within the closed interval of 1 to n, and the predicted data unit is the last element of the output data.

[0119] In some embodiments, the data prediction completion condition includes that the last data unit of the currently predicted data is an end symbol or the length of the currently predicted data reaches the prediction length upper limit value.

[0120] In some embodiments, the currently predicted data further includes a prompt word, which is used to prompt the prediction large model to perform the target prediction task; and / or, the prediction module 34 obtains the prediction result of the target prediction task based on the latest currently predicted data, including: when the currently predicted data does not contain a prompt word, using the latest currently predicted data as the prediction result of the target prediction task; when the currently predicted data contains a prompt word, using the latest currently predicted data as the prediction result of the target prediction task, or, after deleting the prompt word in the latest currently predicted data, using it as the prediction result of the target prediction task.

[0121] In some embodiments, the following steps for training the prediction large model are further included: obtaining original training data, where the original training data includes a plurality of sample data units; dividing the plurality of sample data units into blocks to obtain several data blocks, where the several data blocks include at least one standard data block, and the standard data block includes n sample data units; selecting at least one target data block from the several data blocks, where the at least one target data block includes the standard data block; for each target data block, constructing m - 1 new training data of the target data block, where m is the number of sample data units in the target data block, and the new training data is obtained by replacing the target data block in the original training data with a replacement data block, and the replacement data blocks in the m - 1 new training data are respectively composed of 1 to m - 1 placeholders; training the prediction large model using the original training data and the new training data of each target data block.

[0122] In some embodiments, the original training data and the new training data include sample prompt words, which are used to prompt the prediction large model to perform sample prediction tasks based on the original training data or the new training data; and / or, all data blocks except the last data block among the several data blocks are standard data blocks, and the last data block is a standard data block or the number of sample data units is less than n; and / or, selecting at least one target data block from the several data blocks includes: taking all data blocks as target data blocks.

[0123] Please refer to Figure 4 , Figure 4 It is a schematic framework diagram of an embodiment of a prediction large model training device provided by this application. In this embodiment, the prediction large model training device 40 includes an acquisition module 41, a data chunking module 42, a data selection module 43, and a training module 44. The acquisition module 41 is used to acquire original training data, where the original training data includes a plurality of sample data units; the data chunking module 42 is used to chunk the plurality of sample data units to obtain a number of data chunks, where the number of data chunks includes at least one standard data chunk, and the standard data chunk includes n sample data units; the data selection module 43 is used to select at least one target data chunk from the number of data chunks, where the at least one target data chunk includes a standard data chunk; the data construction module is used to construct m - 1 new training data for each target data chunk, where m is the number of sample data units in the target data chunk, and the new training data is obtained by replacing the target data chunk in the original training data with a replacement data chunk, and the replacement data chunks in the m - 1 new training data are respectively composed of 1 to m - 1 placeholders; the training module 44 is used to train the prediction large model by using the original training data and the new training data of each target data chunk.

[0124] Please refer to Figure 5 , Figure 5 It is a schematic framework diagram of an embodiment of an electronic device provided by this application. In this embodiment, the electronic device 50 includes a memory 51 and a processor 52 that are coupled to each other.

[0125] The memory 51 stores program instructions, and the processor 52 is used to execute the program instructions stored in the memory 51 to implement the steps of any of the above method embodiments. In a specific implementation scenario, the electronic device 50 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 50 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited here.

[0126] Specifically, the processor 52 is used to control itself and the memory 51 to implement the steps of any of the above embodiments. The processor 52 can also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip with signal processing capabilities. The processor 52 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Additionally, the processor 52 can be implemented jointly by integrated circuit chips.

[0127] Please refer to Figure 6 , Figure 6 is a schematic framework diagram of the computer-readable storage medium provided by the present application. The computer-readable storage medium 60 of the embodiments of the present application stores program instructions 61, and when the program instructions 61 are executed, the methods provided by any one of the above methods and any non-conflicting combinations are implemented. Among them, the program instructions 61 can form a program file and be stored in the above computer-readable storage medium 60 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium 60 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0128] In the above solution, the number of input data for the prediction large model each time is n, and each piece of input data includes the position indicating a data unit to be predicted this time, and the positions indicated by each piece of input data are different. Therefore, when using the prediction large model to perform the target prediction task based on each piece of input data respectively, n predicted data units at different positions can be predicted based on each piece of input data. Compared with the existing method where the prediction large model only outputs one predicted data unit each time, the method of constructing the n pieces of input data this time in the present application so that the prediction large model outputs n predicted data units at different positions each time can significantly improve the efficiency of data prediction.

[0129] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0130] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0131] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division manners. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0132] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0134] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0135] The above is only the embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.< / eos> < / m> < / m> < / m> < / m> < / m> < / eos> < / m> < / m> < / m> < / m> < / eos> < / m> < / m> < / m> < / eos> < / m> < / m> < / eos> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / eos> < / eos> < / m> < / m> < / m> < / m> < / m> < / eos> < / m> < / m> < / m> < / m> < / eos> < / m> < / m> < / m> < / eos> < / m> < / m> < / eos> < / m> < / eos> < / eos> < / eos> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / m> < / eos>

Claims

1. A text prediction method, characterized in that: include: Constructing n input texts of this time, wherein each of the input texts includes the current predicted text and indicates the position of a character to be predicted this time, the current predicted text includes the characters predicted by the prediction model when performing the text prediction task before this time, the position of the character to be predicted indicated by each input text is different, and n is greater than one; Utilizing the large prediction model to perform the text prediction task in parallel on each of the input texts to obtain n predicted characters; Adding the n predicted characters to the current predicted text to obtain an updated current predicted text, wherein the position of each predicted character in the current predicted text matches the position indicated by the corresponding input text; The above steps are repeated until the text prediction completion condition is met, and the target predicted text of the text prediction task is obtained based on the latest currently predicted text.

2. The method according to claim 1, characterized in that Among the n input texts, the i-th input text also includes i-1 placeholders, where i is an integer in a closed interval between 1 and n; the position indicated by each input text is after the last element in the input text.

3. The method according to claim 1, characterized in that The method of using the prediction large model to perform the text prediction task on each of the input texts in parallel to obtain n predicted characters includes: The text prediction task is performed in parallel on each of the input texts using the prediction large model to obtain n output texts corresponding to the n input texts respectively, wherein each of the output texts includes a predicted character corresponding to the position indicated by the input text.

4. The method according to claim 3, characterized in that Among the n output texts, the jth output text includes the current predicted text, j-1 placeholders, and a predicted character, wherein j is an integer in a closed interval between 1 and n, and the predicted character is the last element of the output text.

5. The method according to claim 1, characterized in that The text prediction completion condition includes that the last character of the currently predicted text is a terminator or the length of the currently predicted text reaches an upper limit of the predicted length.

6. The method according to claim 1, characterized in that The currently predicted text also includes a prompt word, and the prompt word is used to prompt the prediction model to perform the text prediction task; And / or, obtaining the target predicted text of the text prediction task based on the latest currently predicted text includes: In the case that the current predicted text does not contain a prompt word, taking the latest current predicted text as the target predicted text of the text prediction task; In the case where the current predicted text includes a prompt word, the latest current predicted text is used as the target predicted text of the text prediction task, or the prompt word in the latest current predicted text is deleted and used as the target predicted text of the text prediction task.

7. The method according to claim 1, characterized in that The method also includes the following steps of training the prediction model: Acquire an original training text, wherein the original training text includes a plurality of sample characters; Dividing the plurality of sample characters into blocks to obtain a plurality of text blocks, wherein the plurality of text blocks include at least one standard text block, and the standard text block includes n of the sample characters; Selecting at least one target text block from the plurality of text blocks, wherein the at least one target text block includes the standard text block; For each target text block, construct m-1 new training texts of the target text block, wherein m is the number of sample characters in the target text block, m is greater than 1 and less than or equal to n, and the new training text is obtained by replacing the target text block in the original training text with a replacement text block, and the replacement text blocks in the m-1 new training texts are respectively composed of 1 to m-1 placeholders; The prediction large model is trained using the original training text and the new training text of each target text block.

8. The method according to claim 7, characterized in that The original training text and the new training text include sample prompt words, and the sample prompt words are used to prompt the prediction model to perform a text prediction task based on the original training text or the new training text; And / or, the text blocks except the last text block among the plurality of text blocks are all the standard text blocks, the last text block is the standard text block or the number of the sample characters is less than n; And / or, selecting at least one target text block from the plurality of text blocks comprises: All of the text blocks are taken as the target text blocks.

9. A method for training a large prediction model, characterized in that: include: Acquire an original training text, wherein the original training text includes a plurality of sample characters; Dividing the plurality of sample characters into blocks to obtain a plurality of text blocks, wherein the plurality of text blocks include at least one standard text block, and the standard text block includes n of the sample characters; Selecting at least one target text block from the plurality of text blocks, wherein the at least one target text block includes the standard text block; For each target text block, construct m-1 new training texts of the target text block, wherein m is the number of sample characters in the target text block, m is greater than 1 and less than or equal to n, and the new training text is obtained by replacing the target text block in the original training text with a replacement text block, and the replacement text blocks in the m-1 new training texts are respectively composed of 1 to m-1 placeholders; The prediction large model is trained using the original training text and the new training text of each target text block.

10. A text prediction device, characterized in that: include: A construction module, used to construct the n input texts of this time, wherein each of the input texts includes the current predicted text and indicates the position of a character to be predicted this time, the current predicted text includes the characters predicted by the prediction model when performing the text prediction task before this time, the position indicated by each input text is different, and n is greater than one; A task execution module, used for executing the text prediction task of this time in parallel on each of the input texts using the prediction big model to obtain n predicted characters; A data updating module, configured to add the n predicted characters to the current predicted text to obtain an updated current predicted text, wherein the position of each predicted character in the current predicted text matches the position indicated by the corresponding input text; The prediction module is used to repeat the above steps until the text prediction completion condition is met, and obtain the target predicted text of the text prediction task based on the latest currently predicted text.

11. A prediction large model training device, characterized in that: include: An acquisition module, used for acquiring an original training text, wherein the original training text includes a plurality of sample characters; A data block division module, used for dividing the plurality of sample characters into blocks to obtain a plurality of text blocks, wherein the plurality of text blocks include at least one standard text block, and the standard text block includes n of the sample characters; A data selection module, used for selecting at least one target text block from the plurality of text blocks, wherein the at least one target text block includes the standard text block; A data construction module is used to construct m-1 new training texts of the target text block for each target text block, wherein m is the number of sample characters in the target text block, m is greater than 1 and less than or equal to n, and the new training text is obtained by replacing the target text block in the original training text with a replacement text block, and the replacement text blocks in the m-1 new training texts are respectively composed of 1 to m-1 placeholders; A training module is used to train the prediction large model using the original training text and the new training text of each target text block.

12. An electronic device, characterized in that: comprising a memory and a processor coupled to each other, The memory stores program instructions; The processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions that can be run by a processor, and the program instructions can be executed by the processor to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method, electronic device, and storage medium for training text generation model

    US20210374359A1