Text generation method and system and model training method and system
By generating word elements in parallel and using the diffusion model of bidirectional text information, the problems of low efficiency and poor quality of text generation in the prior art are solved, and efficient and high-quality text generation is achieved.
Patent Information
- Application Number
- CN202510165507.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
Existing autoregressive models are less efficient in text generation, especially when generating long texts, which significantly increases in computational cost and time consumption, while affecting the quality of generated texts.
By generating word elements in parallel during text generation, and using the bidirectional text information of word elements as reference, a diffusion model is used to make multiple rounds of predictions on noise text to generate the target text.
It improves text generation efficiency, while ensuring the quality of text generation, and can generate high-quality text based on the perceived text context.
Smart Images

Figure CN120104737A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence, and in particular to a text generation method and system, and a model training method and system. Background Art
[0002] With the development of science and technology, natural language processing (NLP) technology has gradually demonstrated powerful text generation capabilities and has been widely used in a variety of scenarios that require text generation. For example, in the question-and-answer scenario, natural language processing technology can answer various questions raised by users.
[0003] Although natural language processing technology has achieved good results in text processing and generation, there are still challenges. For example, the existing autoregressive model can gradually generate the next word based on the generated text by predicting word by word. This word-by-word generation method will make the text generation efficiency low, especially when generating long texts, the computational cost and time consumption will increase significantly. At the same time, this method of text generation, which is highly dependent on the generated text, may also affect the quality of the generated text.
[0004] Therefore, it is necessary to provide a text generation method that can improve text generation efficiency while maintaining text generation quality.
[0005] The content of the background technology section is only the information known to the inventor personally, and does not mean that the above information has entered the public domain before the application date of this disclosure, nor does it mean that it can become the prior art of the present disclosure. Summary of the invention
[0006] The present specification provides a text generation method and system, a model training method and system, which generate word units in parallel during the text generation process and use the bidirectional text information of the word units as a reference when generating word units, thereby improving the text generation efficiency while ensuring the quality of text generation.
[0007] In a first aspect, this specification provides a text generation method. The text generation method includes the following steps: obtaining a target prompt text, where the target prompt text represents the requirements for text generation; inputting the target prompt text into a trained diffusion model to generate a target text through the diffusion model, where the target text includes a plurality of tokens, and the target model is trained to have the ability to perform multiple rounds of prediction on a noisy text based on the target prompt text to gradually generate each token in the target text. Among them, in each round of prediction, the diffusion model generates a part of the tokens in the target text in a parallel manner. For a target token in the target text, when generating the target token, the diffusion model uses the forward text information and backward text information of the target token as a reference. The forward text information includes the text information before the target token, and the backward text information includes the text information after the target token.
[0008] In some embodiments, the process of the diffusion model generating the target text includes: generating an initial noisy text, where the initial noisy text includes a plurality of masks; and performing T rounds of iteration based on the initial noisy text, where T is an integer greater than 1. Among them, the process of the i-th round of iteration includes: under the guidance of the target prompt text, performing token prediction on each mask in the current noisy text to obtain a prediction text. When i = 1, the current noisy text is the initial noisy text. When i > 1, the current noisy text is the noisy text updated in the (i - 1)-th round of iteration. In the case of i = T, the target text is determined based on the prediction text. In the case of i < T, a part of the tokens in the prediction text are replaced with masks to obtain an updated noisy text.
[0009] In some embodiments, the step of replacing a part of the tokens in the prediction text with masks to obtain an updated noisy text includes: determining the mask ratio corresponding to the i-th round of iteration, where, in the order of i from 1 to T, the mask ratio gradually decreases; and replacing a part of the tokens in the prediction text that meet the mask ratio with masks to obtain the updated noisy text.
[0010] In some embodiments, the tokens predicted and generated in the current round of iteration in the prediction text form a first token set, and the tokens replaced with masks in the prediction text form a second token set, and the second token set is a subset of the first token set.
[0011] In some embodiments, determining the mask ratio corresponding to the i-th round of iteration includes: determining a target step length based on the value of T; and performing one of the following based on the target step length to obtain the mask ratio corresponding to the i-th round of iteration: when i=1, subtracting the target step length from 1, or when i>1, subtracting the target step length from the mask ratio corresponding to the i-1th round of iteration.
[0012] In some embodiments, the method further includes: obtaining the value of T from configuration information of the user for the diffusion model.
[0013] In some embodiments, replacing some of the words in the predicted text that meet the mask ratio with masks to obtain the updated noisy text includes: using a random strategy to determine a number of words in the predicted text to replace with masks to obtain the updated noisy text, wherein the proportion of the words replaced with masks in the predicted text meets the mask ratio.
[0014] In some embodiments, the process of the i-th iteration further includes: for each word element in the predicted text that is generated in the current iteration and is not replaced by a mask, determining that the word element is generated in the i-th iteration; the target text is divided into a plurality of text segments, each of which includes a plurality of word elements, and the generation order of the word elements in the plurality of text segments satisfies: if the first word element in the j-th text segment is in the t-th text segment, j In the j+1th text segment, any second word in the tth text segment is generated in the tth j Round iteration or t j The iteration after the round iteration is generated, wherein j is an integer greater than or equal to 1, and t j is an integer greater than or equal to 1 and less than or equal to the T.
[0015] In some embodiments, the generating of the initial noisy text includes: generating a noise text composed of a plurality of masks; and splicing the noise text behind the target prompt text to obtain the initial noisy text; when performing word prediction on each mask in the current noisy text, at least part of the content in the target prompt text is used as the forward text information of each mask.
[0016] In some embodiments, the diffusion model adopts a Transformer Encoder network architecture and uses a bidirectional attention mechanism in each round of prediction.
[0017] On the second aspect, this specification also provides a model training method. The model training method comprises the following steps: obtaining a sample data set, wherein each training sample in the sample data set comprises at least one text; obtaining a diffusion model to be trained; and performing multiple rounds of iterative training on the diffusion model using the sample data set, wherein each round of iteration comprises the following steps: obtaining a target sample text from the sample data set, randomly generating a mask ratio within the range of [0,1], replacing word units in the target sample text that meet the mask ratio with masks to obtain noisy text, performing word unit prediction on each mask in the noisy text using the diffusion model to obtain a predicted text, wherein when performing word unit prediction, the diffusion model predicts the masks in the noisy text in a parallel manner, and for any target mask in the noisy text, when the diffusion model performs word unit prediction on the target mask, the forward text information and backward text information of the target mask are used as references, wherein the forward text information comprises text information located before the target mask, and the backward text information comprises text information located after the target mask, and the parameters of the diffusion model are updated with the training goal of minimizing the difference between the predicted text and the sample text.
[0018] In some embodiments, the sample data set includes a first data subset and a second data subset, and the diffusion model is subjected to multiple rounds of iterative training using the sample data set, including: pre-training the diffusion model using the first data subset so that the diffusion model has the ability to recover clean text from noisy texts with different noise content ratios; and fine-tuning the diffusion model using the second data subset so that the diffusion model has the ability to generate text under the guidance of prompt text.
[0019] In some embodiments, each training sample in the first data subset includes a text; during the pre-training process, obtaining the target sample text from the sample data set includes: determining a target training sample from the first data subset, and using the text in the target training sample as the target sample text.
[0020] In some embodiments, each training sample in the second data subset includes a first text and a second text, the first text is used to describe the requirements of text generation, and the second text is a text that meets the requirements; in the process of the fine-tuning training, obtaining the target sample text from the sample data set includes: determining a target training sample from the second data subset, and concatenating the first text and the second text in the target training sample to obtain the target sample text, wherein, when performing word unit prediction on the target mask, at least part of the content in the first text is used as the forward text information of the target mask.
[0021] In some embodiments, the word-grams replaced by masks in the target sample text are word-grams from the second text.
[0022] In some embodiments, the diffusion model adopts a Transformer Encoder network architecture and a bidirectional attention mechanism in the process of word prediction.
[0023] In a third aspect, the specification also provides a text generation system. The text generation system includes: at least one storage medium storing at least one instruction set; and at least one processor communicating with the at least one storage medium, wherein when the text generation system is running, the at least one processor reads the at least one instruction set and executes the text generation method described in the first aspect according to the instructions of the at least one instruction set.
[0024] In a fourth aspect, the present specification also provides a training system. The training system includes: at least one storage medium storing at least one instruction set; and at least one processor communicating with the at least one storage medium, wherein when the training system is running, the at least one processor reads the at least one instruction set and executes the model training method described in the second aspect according to the instructions of the at least one instruction set.
[0025] Other functions of the text generation method and system, model training method and system provided in this specification will be partially listed in the following description. The creative aspects of the text generation method and system, model training method and system provided in this specification can be fully explained by practicing or using the methods, devices and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 The application scenario of the text generation system provided according to the embodiment of this specification is shown;
[0028] Figure 2 shows a system hardware structure diagram provided according to an embodiment of this specification;
[0029] Figure 3 A flowchart of a method for training a diffusion model provided according to an embodiment of this specification is shown;
[0030] Figure 4 A schematic diagram showing a partial training process in pre-training provided according to an embodiment of this specification is shown;
[0031] Figure 5 A schematic diagram showing a partial training process in fine-tuning training provided according to an embodiment of this specification;
[0032] Figure 6 A flowchart of a text generation method provided according to an embodiment of this specification is shown;
[0033] Fig. 7A A schematic diagram showing a process of generating a target text according to an embodiment of this specification;
[0034] Figure 7B A schematic diagram showing a process of generating a target text according to an embodiment of this specification;
[0035] Figure 7C A schematic diagram showing a process of generating a target text according to an embodiment of this specification; and
[0036] Figure 8 A diagram showing a generation process of multiple word-grams in a target text provided according to an embodiment of the present specification is shown. DETAILED DESCRIPTION
[0037] The following description provides specific application scenarios and requirements of this specification, with the purpose of enabling those skilled in the art to make and use the contents of this specification. Various local modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but to the widest scope consistent with the claims.
[0038] The terms used herein are only used for the purpose of describing specific example embodiments and are not restrictive. For example, unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an" and "the" may also include plural forms. When used in this specification, the terms "include", "comprise" and / or "contain" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups or that other features, integers, steps, operations, elements, components and / or groups may be added in the system / method.
[0039] In view of the following description, these and other features of the present specification, as well as the operation and function of the related elements of the structure, and the economy of the combination and manufacture of the parts can be significantly improved. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0040] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0041] For the convenience of description, the terms involved in the specification are first explained as follows:
[0042] Diffusion models are a generative model based on deep learning. They generate data by simulating the diffusion process of data from order to disorder, and the reverse diffusion process from disorder to order. In text generation, diffusion models generate new text by gradually adding noise to the text and then learning how to gradually remove the noise.
[0043] A token can be understood as the smallest meaningful unit in text data. It can be a word, number, symbol, or other string fragments that can be processed individually. A token is the basic unit for language models to process text and is used to convert natural language into a form that can be processed by machines.
[0044] Before describing the specific embodiments of this specification, the application scenarios to which the technical solutions provided in this specification are applicable are first introduced.
[0045] The text generation technology in natural language processing technology can generate various types of texts according to the prompts, instructions or contexts input by the user, so it has a wide range of applications in many fields and scenarios. For example, text generation technology can be used for content creation, such as generating news text, assisting in the creation of literary works, and so on. For another example, text generation technology can be used for text analysis, such as judging the emotional tendency of the text, text management and classification, and so on. For another example, text generation technology can be used in question-and-answer or dialogue systems, such as generating accurate answers based on user questions. Among them, under the question-and-answer system, text generation technology encounters more challenges. Text generation technology needs to understand the questions raised by users or systems containing complex contexts, learn knowledge from massive texts, and thus generate fluent, coherent, accurate and in-depth answers in a timely manner. For the sake of ease of narration, this application mainly introduces the application of text generation technology in question-and-answer or dialogue systems as an example.
[0046] In the prior art, an autoregressive language model can be used for text generation to achieve human-computer dialogue. The autoregressive language model generates the entire text by gradually generating the next word based on the generated context in a sequential prediction manner. Since the autoregressive model is based on the conditional probability chain rule, it gradually generates the probability distribution of each word and depends on the results of the previous text generation, so the autoregressive model must generate the next word in sequence. This limits the rate of text generation, especially when generating long texts, this serialization will significantly increase the computing device overhead. In addition, the autoregressive language model uses a unidirectional attention mechanism, which means that the autoregressive model cannot perceive the context that may be related to the text when predicting text, which affects the generation results of the model.
[0047] In the prior art, bidirectional encoder representation (Bidirectional Encoder Representation from Transformers, BERT) is also used for text generation. However, due to BERT's training mechanism, the mask ratio of BERT during the training phase is fixed, usually 15%, so that BERT mainly understands the semantics and background of the question based on contextual information. Therefore, BERT can usually only complete questions similar to cloze, and has difficulty in achieving smooth human-computer dialogue.
[0048] This specification provides a text generation method and system, and a model training method and system. When generating text, the diffusion model can perform multiple rounds of predictions on the text containing masks of any proportion based on the target prompt text, thereby generating the target text. In each round of prediction, the diffusion model can generate word units in the target text in parallel, thereby improving the efficiency of text generation. Moreover, when generating any word unit in the target text, the diffusion model can use the forward text information and backward text information of the word unit as reference at the same time, so that when generating text, it can perceive the context related to the text and ensure the quality of text generation.
[0049] Figure 1 FIG. 1 shows an application scenario of a text generation system provided according to an embodiment of this specification. Figure 1 As shown, scenario 100 may include a model training system 110 and a text generation system 120 .
[0050] Scenario 100 can be divided into two stages, the model training stage and the text generation stage.
[0051] In the model training phase, the model training system 110 can perform multiple rounds of iterative training on the diffusion model based on the sample data set. After the training is completed, the model training system 110 can deploy the trained diffusion model in the text generation system 120. The diffusion model can adopt the Transformer Encoder network architecture and adopt a bidirectional attention mechanism in the iterative process.
[0052] In the text generation stage, the text generation system 120 can call the trained diffusion model and perform multiple rounds of predictions on the noisy text based on the target prompt text to generate the target text. Among them, the target prompt text can represent the needs of text generation. The target text can be a text that meets the text generation needs. The target prompt text can be provided by the user or automatically generated by the system. The text generation needs can be the specific requirements and expectations of the user or system for the generated text. For example, the needs can be guiding information about content, style, format, length, etc., which is used to guide the model to generate text that meets specific goals. For example, the target prompt text can be a prompt input by the user or the system. Prompt refers to a prompt word, which is an input guide or instruction. Prompt can be understood as a language for communicating with the model to tell the model what it wants or intends to generate.
[0053] In some embodiments, the model training system 110 may store data and instructions for implementing a training method for a diffusion model, and may execute or be used to execute the data and instructions. In some embodiments, the model training system 110 may include a hardware device with a data information processing function and a necessary program to drive the hardware device to work.
[0054] In some embodiments, the text generation system 120 may store data and instructions for implementing the text generation method, and may execute or be used to execute the data and instructions. In some embodiments, the text generation system 120 may include a hardware device with a data information processing function and a necessary program required to drive the hardware device to work.
[0055] It is understandable that the model training system 110 and the text generation system 120 may correspond to the same system or to different systems, and this specification does not impose any limitation on this.
[0056] It should be noted that the model training system 110 can correspond to one device or a device cluster, and this specification does not limit this. When the model training system 110 corresponds to one device, the training method of the diffusion model can be completely executed on the device. When the model training system 110 corresponds to a device cluster, the training method of the diffusion model can be executed in coordination on multiple devices corresponding to the device cluster, or the training method of the diffusion model can also have other execution methods, and this specification does not limit this.
[0057] The text generation system 120 may correspond to one device or to a device cluster, and this specification does not limit this. When the text generation system 120 corresponds to one device, the text generation method may be executed entirely on the device. When the text generation system 120 corresponds to a device cluster, the text generation method may be executed in coordination on multiple devices corresponding to the device cluster, or the text generation method may have other execution modes, and this specification does not limit this.
[0058] In some embodiments, the scenario 100 may also include a user 130 and a terminal 140 .
[0059] The user 130 may be any user who generates text.
[0060] The terminal 140 may be a device for interacting with the user 130. The user 130 may input various data or instructions to the terminal 140 through the input and output devices of the terminal 140, and receive data or results from the terminal 140. In some embodiments, the terminal 140 may include a hardware device with a data information processing function and an application necessary to drive the hardware device to work. The application can provide the user 130 with the ability and interface to interact with the outside world through the network.
[0061] Terminal 140 may receive a target prompt text input by user 130, and output the target text to user 130. The target prompt text may represent the text generation requirement of user 130. For example, terminal 140 may receive <What is the previous sentence of "not sticking to one pattern to select talents?" Just output the sentence> input by user 130, and output <I advise the heaven to shake itself up again.> to user 130. For another example, terminal 140 may receive <What is the previous sentence of "not sticking to one pattern to select talents" input by user 130? Just output the sentence><Please help me translate intoChinese:‘What is now proved was once only imagined’> , and outputs <what is now proven to be just imagination> to the user 130. For another example, the terminal 140 may receive continuous questions from the user 130 and output coherent answers to the user 130, thereby achieving a smooth and natural dialogue with the user 130.
[0062] In some embodiments, the terminal 140 may include a mobile device, a tablet computer, a laptop computer, a built-in device of a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device, or the like, or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, or the like, or any combination thereof. In some embodiments, the smart mobile device may include a smart phone, a personal digital assistant, a gaming device, a navigation device, or the like, or any combination thereof. In some embodiments, the built-in device in the motor vehicle may include an onboard computer, an onboard TV, or the like.
[0063] As an example, when user 130 needs to generate text, he can send the target prompt text and text generation instructions to the text generation system 120 through terminal 140. The text generation system 120 receives the target prompt text and responds to the text generation instruction to execute the text generation method provided in this specification.
[0064] It should be understood that Figure 1 The numbers of the model training system 110, the text generation system 120, the users 130, and the terminals 140 shown in the figure are merely illustrative and can be any number according to the implementation requirements.
[0065] Figure 2 A hardware structure diagram of a system 200 provided according to an embodiment of the present specification is shown.
[0066] like Figure 2 As shown, the system 200 may be Figure 1 The model training system 110 or the text generation system 120 in.
[0067] The system 200 includes at least one storage medium 230 and at least one processor 220. In some embodiments, the system 200 may further include a communication port 250 and an internal communication bus 210. In addition, the system 200 may further include an I / O component 260.
[0068] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250 and the I / O component 260.
[0069] I / O component 260 supports input / output between the system 200 and other components.
[0070] The communication port 250 is used for data communication between the system 200 and the outside world. For example, the communication port 250 can be used for data communication between the system 200 and a network. The communication port 250 can be a wired communication port or a wireless communication port.
[0071] In some embodiments, the network may be any type of wired or wireless network, or a combination thereof. For example, the network may include a cable network, a wired network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network TM, a short-range wireless network (ZigBee TM), a near field communication (NFC) network, or a similar network.
[0072] In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station or an Internet exchange point. Through the access point, one or more components of each device corresponding to the system 200 can be connected to the network to exchange data or information.
[0073] The storage medium 230 may include a data storage device. The data storage device may be a non-temporary storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. The storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may include a computer program code, and the computer program code may include a program, a routine, an object, a component, a data structure, a process, a module, etc. for executing the text generation method or the diffusion model training method provided in this specification.
[0074] The processor 220 may be in communication with the storage medium 230. The processor 220 is used to execute the at least one instruction set. When the system 200 is running, the processor 220 reads the at least one instruction set and executes the text generation method or the diffusion model training method provided in this specification according to the instructions of the at least one instruction set.
[0075] The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.
[0076] For illustration purposes only, only one processor 220 is shown in the system 200 in the accompanying drawings. However, it should be noted that the system 200 described in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be performed by one processor as described in this specification, or may be performed jointly by multiple processors. For example, if the processor 220 of the system 200 described in this specification performs step A and step B, it should be understood that step A and step B may also be performed jointly or separately by two different processors 220 (e.g., the first processor performs step A, the second processor performs step B, or the first and second processors perform steps A and B together).
[0077] Figure 3 FIG. 1 is a flow chart of a diffusion model training method P300 provided according to an embodiment of the present specification. The model training system 110 may execute the diffusion model training method P300. Figure 3As shown, the training method P300 of the diffusion model includes the following steps.
[0078] S310: Obtain a sample data set, where each training sample in the sample data set includes at least one text.
[0079] The model training system 110 may obtain a sample data set for subsequent training of the diffusion model to be trained. The sample data set may include multiple training samples. Each training sample may include at least one text.
[0080] In some embodiments, the sample data set may include a first data subset and a second data subset. The first data subset and the second data subset may include different types of training samples. For example, the first data subset includes a plurality of training samples, each of which may include a text. For example, each training sample may be an independent text segment without contextual prompts or instructions.
[0081] As an example, the first data subset may include training samples in the following form:
[0082] {Training sample A: [Su Shi, also known by his courtesy name Zizhan and pseudonym Dongpo Jushi, was a writer, calligrapher and painter in the Northern Song Dynasty. His representative works such as "Nian Nujiao·Reminiscence of the Past at Chibi" are widely circulated. Su Shi had a rough life, but he was open-minded and optimistic. His artistic achievements and personal charm have far-reaching influence. ];
[0083] Training sample B: [Shenzhen weather is good, cloudy, temperature 14℃ to 23℃, easterly wind 2-3];
[0084] …
[0085] Training sample X: [The Shawshank Redemption is a classic drama film directed by Frank Darabont. The film tells the story of banker Andy who was wrongly imprisoned. He relied on his wisdom and perseverance in prison to finally achieve self-redemption and regain his freedom]}.
[0086] The second data subset may include multiple training samples, each of which may include a text pair. Each training sample may include a first text and a second text. The first text and the second text correspond to each other. The first text may be used to describe the requirements for text generation. The second text may be a text that meets the aforementioned text generation requirements.
[0087] In some embodiments, the training samples in the second data subset may be a conversation data set. For example, the training samples are data of human-machine or machine-machine conversations. For example, the first text may be a prompt input by a user or a system. The second text may be text output by the model according to the prompt. For another example, the training samples are a person-to-person conversation data set, that is, data of communication between real people.
[0088] As an example, the second data subset may include training samples in the following form:
[0089] {Training sample a:
[0090] [First text: What is the previous sentence of 'but I hear the sound of people'?
[0091] Second text: "But I only hear the sound of people's voices" is preceded by "No one is seen in the empty mountain.";
[0092] Training sample b:
[0093] [First text: Recommend a book to me and give the reason for the recommendation, 30 words.
[0094] Second text: I recommend The Little Prince, which explores friendship, love and human nature with concise words and profound meaning, and is suitable for readers of all ages.]
[0095] …
[0096] Training sample x:
[0097] [First text: Please help me translate the initial two lines of therenowned poem 'TheRoad Not Taken' into Chinese.
[0098] Second text: The first two lines of "The Road Not Taken" by Robert Frost can be translated into Chinese as: "Two roads diverged in a yellow wood, and regrettably I could not take both"]}
[0099] In some embodiments, the sample data set may include training samples in different languages, and the training samples in different languages are trained together to improve the cross-language understanding ability of the model. In some embodiments, the different languages in the sample data set allow the training samples to be trained separately to improve the quality of content generation for a specific language. This specification does not limit this.
[0100] In some embodiments, the sample data set obtained by the model training system 110 may be pre-processed data. For example, the text in the sample data set has been cleaned, standardized, etc. to ensure the quality and consistency of the text.
[0101] The sample data set may be obtained from public academic resources, or from a specific industry database, which is not limited in this specification.
[0102] S320: Obtain a diffusion model to be trained.
[0103] The diffusion model is a generative language model that can generate related text based on the input sentence.
[0104] In some embodiments, the diffusion model can adopt the Transformer Encoder network architecture and adopt a bidirectional attention (Bi-Attention) mechanism in each round of prediction. Among them, Transformer Encoder is an encoder network based on the Transformer (a language model using an attention mechanism) architecture. In Transformer Encoder, its self-attention mechanism (Self-Attention) is bidirectional. This means that when calculating the representation of each word in the text, the self-attention mechanism of Transformer Encoder allows each word to simultaneously consider the information of all text words before and after the word. This consideration of bidirectional contextual information enables the model to understand the relationship between words more comprehensively, thereby generating more accurate contextual representations.
[0105] And because the self-attention mechanism is bidirectional, Transformer Encoder can process the entire text sequence in parallel, without having to process the words in the text one by one in order. In addition, because the diffusion model can also denoise multiple words in parallel during the reverse generation process (the process of denoising the text), the diffusion model using the Transformer Encoder network structure can realize text generation using the bidirectional attention mechanism after training, and each round of prediction process in the generated text can generate multiple words in parallel. The specific text generation process will be introduced later.
[0106] S330: Perform multiple rounds of iterative training on the diffusion model using the sample data set.
[0107] In each round of iterative training, the model training system 110 uses at least one training sample in the sample data set to perform forward propagation on the diffusion model, calculate the loss, perform back propagation, and update the parameters, thereby completing a round of iterative training.
[0108] Please continue to see Figure 3 , each iterative training process includes steps S331-S334.
[0109] S331: Obtain target sample text from the sample data set.
[0110] S332: randomly generate a mask ratio in the range of [0, 1], and replace the word units in the target sample text that meet the mask ratio with the mask to obtain a noisy text.
[0111] S333: The diffusion model is used to perform word-unit prediction on each mask in the noisy text to obtain a predicted text, wherein, when performing word-unit prediction, the diffusion model predicts the masks in the noisy text in parallel, and for any target mask in the noisy text, when the diffusion model performs word-unit prediction on the target mask, the forward text information and backward text information of the target mask are used as references, the forward text information includes text information located before the target mask, and the backward text information includes text information located after the target mask.
[0112] S334: Taking minimizing the difference between the predicted text and the sample text as a training goal, the parameters of the diffusion model are updated.
[0113] As mentioned above, the sample data set may include a first data subset and a second data subset. The first data subset and the second data subset may be used for different training processes of the model respectively. In some embodiments, the performing multiple rounds of iterative training on the diffusion model using the sample data set may include:
[0114] (1) Pre-training process: The diffusion model is pre-trained using the first data subset so that the diffusion model has the ability to recover clean text from noisy texts with different noise ratios.
[0115] Pre-training can learn the common features of language through large-scale general data and improve the generalization ability of the model.
[0116] For example, the diffusion model can be pre-trained on a first data subset containing tens of thousands of words. During the pre-training process, the model training system 110 can use the first data subset to perform multiple rounds of iterative training on the diffusion model, so that the diffusion model can remove noise in the text and restore clear and accurate text content.
[0117] Figure 4 A schematic diagram of a part of the training process in the pre-training provided in the embodiment of this specification is shown. Figure 4 , taking one iterative training process in multiple rounds of iterative training under pre-training as an example to introduce. Other iterative processes are similar and will not be described in detail.
[0118] (i) Obtaining target sample text from the sample data set.
[0119] In some embodiments, during the pre-training process, obtaining the target sample text from the sample data set may include: determining a target training sample from the first data subset, and using the text in the target training sample as the target sample text.
[0120] The target training sample of pre-training may include at least one training sample in the first data subset. For example, the target training sample may include at least one text.
[0121] In some embodiments, in order to efficiently utilize computing resources and improve training speed, the model training system 110 can divide the sample data set into multiple batches. Each batch includes multiple training samples. Each time a batch is processed, the model performs a forward propagation, calculates the loss, performs a back propagation, and updates the parameters. Therefore, the target training sample can include multiple training samples, that is, multiple texts.
[0122] For ease of presentation, Figure 4 Only two training samples, two texts, in the target training samples are shown.
[0123] The target training samples of different rounds of iterative training may be the same samples or completely different samples, which is not limited in this specification.
[0124] In some embodiments, determining the target training sample from the first data subset may be determining training samples with higher text quality and better text diversity from the first data subset as target training samples to improve model training efficiency and enhance model performance.
[0125] (ii) randomly generating a mask ratio in the range of [0, 1], and replacing the word units in the target sample text that meet the mask ratio with the mask to obtain a noisy text.
[0126] The target sample text may include multiple word units. After randomly generating the mask ratio, the model training system 110 randomly selects word units of the mask ratio from the multiple word units for masking, for example, replacing the randomly selected word units with special mask marks, thereby obtaining noisy text.
[0127] For example, the target sample text includes a text, and the text content is: [The weather is very good today]. The word units included in the text may be [ / today / , / weather / , / very / , / good / ]. When the mask ratio is 50%, the model training system 110 can randomly select 50% of the word units from the four word units for masking. For example, if [ / today / , / very / ] is randomly selected, these two word units will be replaced with a special mask mark [ / MASK / ]. The word unit sequence of the noisy text after replacement is: [ / MASK / , / weather / , / MASK / , / good / ].
[0128] The mask ratios of different rounds of iterative training are randomly generated and can be the same or different, and this specification does not limit this.
[0129] For example, Figure 4 The first text shown in includes 8 tokens T1 to T8, with a randomly generated mask ratio of 25%. Tokens T2 and T7 are replaced with mask M, and a noisy text is obtained. The second text includes 10 tokens T1 to T10, with a randomly generated mask ratio of 40%, tokens T1, T4, T7, and T8 are replaced with mask M, and a noisy text is obtained.
[0130] (iii) performing word-unit prediction on each mask in the noisy text by the diffusion model to obtain a predicted text. When performing word-unit prediction, the diffusion model predicts the masks in the noisy text in parallel, and for any target mask in the noisy text, when the diffusion model performs word-unit prediction on the target mask, the forward text information and backward text information of the target mask are used as references, wherein the forward text information includes text information located before the target mask, and the backward text information includes text information located after the target mask.
[0131] The noisy text may be a text containing a mask of any proportion (in the range of 0-1). The mask may represent noise introduced by reasons including but not limited to misspellings, grammatical errors, random characters or confusing words in the text.
[0132] The diffusion model can predict masks in noisy text in a parallel manner. For example, the diffusion model can simultaneously predict the word units corresponding to all masks in the noisy text.
[0133] The target mask can be any mask in the noisy text. When predicting the target mask in the noisy text, the diffusion model can use the bidirectional text information of the target mask as a reference to perform word unit prediction. The diffusion model can use the text information before the target mask and the text information after the target mask as a reference to perform word unit prediction on the target mask.
[0134] The text information may be information about unmasked word units, or the relationship between unmasked word units, or text structure information formed by unmasked word units, or text structure information between unmasked word units and masks, and so on.
[0135] For example, Figure 4Taking the first noisy text in as an example, the diffusion model can perform word-meta prediction on the two masks in the noisy text at the same time. When predicting the mask at word-meta T2, the diffusion model can use the text information before T2—word-meta T1, and the text information after T2—word-meta T3 to T6, T8 and the relationship between word-meta, etc., to perform word-meta prediction on the mask M at T2. In some embodiments, since multiple masks are predicted for word-meta at the same time, the diffusion model can also use the text information that may be predicted for masks at other positions as a reference when predicting the mask at a certain position. For example, when predicting the mask at word-meta T2, the backward text information of word-meta T2 can also include information that may be predicted at the mask of T7. The diffusion model can use the information that may be predicted by the mask at T7 as a reference to perform mask prediction to obtain word-meta T2'.
[0136] Similarly, the mask at T7 in the noisy text is predicted to obtain word unit T7', and then the predicted text is obtained.
[0137] Through the diffusion model Figure 4 The second noisy text in is used for word unit prediction to obtain word units T1', T4', T7', and T8', and then the predicted text is obtained.
[0138] (iv) Updating the parameters of the diffusion model with the training objective of minimizing the difference between the predicted text and the sample text.
[0139] The difference between the predicted text and the sample text represents the deviation between the text generated by the diffusion model under the current parameter settings and the real text. By updating the parameters of the diffusion model, this deviation can be gradually reduced, thereby improving the quality of the text generated by the diffusion model.
[0140] Since the training process performs mask replacement according to the randomly generated mask ratio in the range of [0,1], more diverse noisy texts can be obtained after mask replacement, so that the trained diffusion model has the ability to predict words for noisy texts with any noise ratio. Compared with the use of a fixed mask ratio, the diffusion model can recover or generate clean text from noisy texts of different degrees, or even from all-noisy texts, making the diffusion model more robust when facing different types of noisy texts, enhancing the model's ability to process texts of variable lengths, and enabling the diffusion module to generate text based on text generation requirements to achieve smooth question-answering or dialogue functions.
[0141] (2) Fine-tuning training process: The diffusion model is fine-tuned using the second data subset so that the diffusion model has the ability to generate text under the guidance of the prompt text.
[0142] Fine-tuning training can further optimize the model through task-specific data to adapt it to the needs of a specific task, thereby improving task performance.
[0143] For example, the diffusion model can be fine-tuned on the second data subset containing large-scale data. During the fine-tuning training process, the model training system 110 can use, for example, a text pair or a dialogue data set in the form of a first text and a second text to perform multiple rounds of iterative training on the diffusion model, thereby enhancing the model's ability to generate text based on specific text generation requirements and enhancing the model's ability to follow instructions, so as to achieve a fluent dialogue function of the model.
[0144] Figure 5 A schematic diagram of a part of the training process in the fine-tuning training provided in an embodiment of this specification is shown. Figure 5 , taking one iterative training process in multiple rounds of iterative training under fine-tuning training as an example to introduce. Other iterative processes are similar and will not be described in detail.
[0145] (i) Obtain target sample text from the sample dataset.
[0146] In some embodiments, during the fine-tuning training process, obtaining the target sample text from the sample data set includes: determining a target training sample from the second data subset, and concatenating the first text and the second text in the target training sample to obtain the target sample text.
[0147] The target training text of the fine-tuning training may include at least one training sample in the second data subset. For example, the target training sample may include at least one text pair of a first text and its corresponding second text.
[0148] Similar to the pre-training process, the target training samples of the fine-tuning training may also include multiple training samples, that is, multiple text pairs of the first text and the second text.
[0149] For ease of presentation, Figure 5 Only one training sample and one text (text pair) in the target training sample are shown. The text pair may include the concatenated first text and the second text.
[0150] The target training samples of different rounds of iterative training may be the same samples or completely different samples, which is not limited in this specification.
[0151] (ii) randomly generating a mask ratio in the range of [0, 1], and replacing the word units in the target sample text that meet the mask ratio with the mask to obtain a noisy text.
[0152] The first text and the second text in the target sample text may include multiple word-grams. For example, Figure 5 The text pair shown in FIG. 1 includes 10 word units T1 to T10, wherein the first text includes word units T1 to T4, and the second text includes word units T5 to T10.
[0153] The purpose of training the diffusion model is to enable it to generate high-quality text based on text generation requirements. Since the first text represents the text generation requirements, this part of the text is usually directly provided by the user or the system and is known. The diffusion model can obtain the first text in its entirety, and the diffusion model usually does not need to predict and generate the content of the first text. The second text is the text that meets the requirements, and this part of the text is usually predicted and generated by the model.
[0154] Therefore, during fine-tuning training, when the model training system performs masking on the word units in the target sample text, the masking process can be performed only on the second text to train the diffusion model's ability to generate text based on text generation requirements. Therefore, in some embodiments, the word units replaced with masks in the target sample text are word units from the second text. The model training system can replace the word units in the second text that meet the mask ratio with masks to obtain noisy text.
[0155] For example, Figure 5 The random mask ratio in is 40%, the word units T5 and T8 in the second text are replaced by mask M, and a noisy text is obtained.
[0156] The mask ratios of different rounds of iterative training are randomly generated and can be the same or different, and this specification does not limit this.
[0157] (iii) performing word-unit prediction on each mask in the noisy text by the diffusion model to obtain a predicted text. When performing word-unit prediction, the diffusion model predicts the masks in the noisy text in parallel, and for any target mask in the noisy text, when the diffusion model performs word-unit prediction on the target mask, the forward text information and backward text information of the target mask are used as references, wherein the forward text information includes text information located before the target mask, and the backward text information includes text information located after the target mask.
[0158] Due to the existence of the first text, the diffusion model needs to perform word prediction on the target mask based on the text generation requirements represented by the first text, and use the forward text information and backward text information of the target mask as a reference.
[0159] As mentioned above, the first text can be concatenated before the second text, so when performing word prediction on the target mask, at least part of the content in the first text can be used as the forward text information of the target mask. Since there is no mask in the first text, the first text is located before all masks, so the first text can be used as the forward text information of all masks.
[0160] For example, the diffusion model Figure 5 When predicting a word unit with the mask at word unit T5, the mask at T5 can be predicted based on the text information before T5 - word units T1 to T4, and the text information after T5 - word units T6 to T7, T9 to T10 and the relationship between word units.
[0161] Diffusion models can be used to Figure 5 The mask located at the word unit T5 and the mask located at the word unit T8 are used to perform word unit prediction at the same time to obtain word unit T5' and word unit T8', and then obtain the predicted text.
[0162] The specific process can refer to the description of the above pre-training process, which will not be repeated here.
[0163] (iv) Updating the parameters of the diffusion model with the training objective of minimizing the difference between the predicted text and the sample text.
[0164] The specific content can refer to the description of the above pre-training process, which will not be repeated here.
[0165] In some embodiments, in order to improve the performance of the diffusion model, the model parameters of the diffusion model provided in the present application are in the billions, for example, about 7 billion, about 8 billion, about 9 billion, etc. According to the scaling law, as the number of parameters in the model increases, the performance of the model usually improves according to the power law. According to the current technical level, when the model parameters are in the billions, on the one hand, a more ideal model performance can be achieved, and on the other hand, the cost of model training (such as time cost, resource cost, training data cost, etc.) can be controlled within a certain range, thereby achieving a balance between training cost and model performance. With the development of technology, the model parameters of the diffusion model can also be at a higher level, such as tens of billions or even higher.
[0166] The diffusion model trained in this application performs well under multiple model evaluation indicators including Massive Multitask Language Understanding (MMLU) and Chinese Massive Multitask Language Understanding (CMMLU), and can achieve fluent human-computer dialogue.
[0167] Figure 6 1 shows a flowchart of a text generation method P600 provided according to an embodiment of the present specification. The text generation system 120 can execute the text generation method P600. Figure 6 As shown, the text generation method P600 includes the following steps:
[0168] S610: Obtain a target prompt text, where the target prompt text represents a text generation requirement.
[0169] The target prompt text may be input by the user or generated by the system. The target prompt text may refer to the above description and will not be described in detail here.
[0170] S620: Input the target prompt text into a trained diffusion model to generate a target text through the diffusion model, wherein the target text includes multiple word-grams, and the diffusion model is trained to have the ability to perform multiple rounds of predictions on the noisy text based on the target prompt text to gradually generate each word-gram in the target text. In each round of prediction, the diffusion model generates some word-grams in the target text in parallel, and for the target word-gram in the target text, the diffusion model uses the forward text information and backward text information of the target word-gram as reference when generating the target word-gram, wherein the forward text information includes the text information before the target word-gram, and the backward text information includes the text information after the target word-gram.
[0171] When generating text, the diffusion model can perform multiple rounds of predictions on the text containing masks of any proportion based on the target prompt text to generate the target text. In each round of prediction, the diffusion model can generate word units in the target text in parallel, thereby improving the efficiency of text generation. In addition, when generating any word unit in the target text, the diffusion model can simultaneously use the word unit's forward text information and backward text information as references, so that when generating text, it can perceive the context related to the text and ensure the quality of text generation.
[0172] In some embodiments, the diffusion model can adopt a Transformer Encoder network architecture and use a bidirectional attention mechanism in each round of prediction. The specific content can be referred to the above description and will not be repeated here.
[0173] 7A to 7C A schematic diagram of a target text generation process provided according to an embodiment of the present specification is shown. 7A to 7C TX in stands for word unit, and M stands for mask. TX can be T1 to T6. 7A to 7C Describe the process of generating the target text.
[0174] (i) Generate an initial noisy text, wherein the initial noisy text includes multiple masks.
[0175] In some embodiments, generating the initial noisy text may include: generating a noise text composed of multiple masks; and splicing the noise text behind the target prompt text to obtain the initial noisy text; when performing word prediction on each mask in the current noisy text, at least part of the content in the target prompt text is used as the forward text information of each mask.
[0176] The number of masks included in the noise text is related to the number of tokens in the pre-generated target text. Each mask will correspond to a token in the target text after the subsequent prediction process. In some text generation tasks, the target text may be long and contain a large number of tokens. Therefore, the number of masks included in the noise text can be relatively large to ensure that the diffusion model has enough space to smoothly generate long texts.
[0177] In the initial noisy text after concatenation, the target prompt text is located before all the masks. When the diffusion model performs word prediction, at least part of the data in the target prompt text can be used as the forward text information of each mask to ensure the generation of more coherent and consistent text.
[0178] In the subsequent iterative process, the diffusion model gradually predicts word units from multiple masks included in the initial noisy text, thereby obtaining the final target text.
[0179] As an example, Fig. 7A A noisy text including 12 masks is shown in FIG. The noisy text is concatenated with the target prompt text to obtain the initial noisy text.
[0180] (ii) performing T rounds of iterations based on the initial noisy text, where T is an integer greater than 1.
[0181] In some embodiments, the value of T is automatically set by the system. In some embodiments, the text generation system 120 can obtain the value of T from the user's configuration information for the diffusion model. For example, the text generation interface displayed by the terminal (client) may include a setting bar, and the user can set the number of rounds of iterations performed by the diffusion model based on the initial noisy text in the setting bar.
[0182] As an example, 7A to 7C Three of the six iterations of the target text generation process are shown. Fig. 7A It shows the process diagram of the first round of iteration when i=1. Figure 7B The process diagram of the second round of iteration is shown when i=2. Figure 7C The process diagram of the last round of iteration when i=T=6 is shown.
[0183] The i-th iteration process in T rounds of iterations includes:
[0184] (1) Under the guidance of the target prompt text, word units are predicted for each mask in the current noisy text to obtain a predicted text.
[0185] The current noisy text may be a text including a mask of any proportion (within the range of 0-1).
[0186] like Fig. 7A As shown, when i=1, the current noisy text is the initial noisy text. The current noisy text includes 12 masks. The diffusion model needs to perform word unit prediction on these 12 masks to obtain the predicted text.
[0187] like FIG. 7B to FIG. 7C As shown in , when i>1, the current noisy text is the noisy text after the i-1th round of iterative update. Figure 7B As shown in , when i=2, the current noisy text is the noisy text after the first round of iteration update. The current noisy text includes 10 masks. The diffusion model needs to perform word unit prediction on these 10 masks to obtain the predicted text. For another example, Figure 7C As shown, i=T=6, and the current noisy text is the noisy text after the (T-1)th round, that is, the fifth round of iterative update. The current noisy text includes 2 masks. The diffusion model needs to perform word unit prediction on these 2 masks to obtain the predicted text.
[0188] In each round of prediction, the diffusion model can simultaneously and parallelly predict the masks in the current noisy text. For example, the diffusion model can simultaneously predict all masks in the current noisy text. For each mask in the current noisy text, under the guidance of the target prompt text, the word unit prediction of the mask is performed with reference to the forward text information and backward text information corresponding to the mask. When all masks are predicted, the predicted text is obtained.
[0189] by Fig. 7AFor example, the current noisy text includes 12 masks. For each of these 12 masks, the diffusion model uses the background and information provided by the target prompt text and makes a token prediction for the mask with reference to the forward text information and backward text information corresponding to the mask. The diffusion model can predict 12 masks simultaneously to obtain 12 T1s that make up the predicted text (it should be noted that for the convenience of illustration in the attached drawings, the tokens predicted in this round of iteration are all represented as T1. In actual applications, the 12 T1s represent different tokens). For example, when the diffusion model predicts the token corresponding to the 3rd mask, the diffusion model will simultaneously consider the information of the tokens that may be predicted for the two masks before it and the information of the tokens that may be predicted for the 9 masks after it, so as to ensure the coherence of the context content.
[0190] The method of making token predictions using bidirectional text information during the iteration process can refer to the description of the model training process and will not be elaborated here.
[0191] Similarly, in the 2nd round of iteration, the diffusion model can predict 10 masks in the current noisy text simultaneously to obtain 8 T2s in the predicted text. In the 6th round of iteration, the diffusion model can predict 2 masks in the current noisy text simultaneously to obtain 2 T6s in the predicted text.
[0192] (2) In the case of i = T, determine the target text based on the predicted text. In the case of i < T, replace some tokens in the predicted text with masks to obtain an updated noisy text.
[0193] Since i = T is the last round of iteration, the target text can be determined based on the predicted text generated in this round of iteration. For example, use the text content in the predicted text other than the target prompt text as the target text.
[0194] The following mainly introduces the case of i < T.
[0195] In the case of i < T, replace some tokens in the predicted text with masks to obtain an updated noisy text.
[0196] In some embodiments, the step of replacing some tokens in the predicted text with masks to obtain an updated noisy text may include the following steps:
[0197] (i) Determine the mask ratio corresponding to the i-th round of iteration, where, in the order of i from 1 to T, the mask ratio gradually decreases.
[0198] As the number of iteration rounds increases, the mask ratio gradually decreases.
[0199] In some embodiments, the above step (i) may include the following steps: determining a target step size based on the value of T; and performing one of the following (a) and (b) based on the target step size to obtain a mask ratio corresponding to the i-th iteration.
[0200] (a) When i=1, the target step size is subtracted from 1.
[0201] (b) When i>1, the target step size is subtracted from the mask ratio corresponding to the i-1th iteration.
[0202] As mentioned above, the value of T can be set automatically by the system or obtained from the configuration information of the diffusion model set by the user. Since the Tth iteration directly outputs the predicted text as the target text, the Tth iteration does not perform the operation of replacing word units with masks, and the mask ratio can be considered to be 0.
[0203] As an example, T=4, the diffusion model performs 4 iterations, and the target step size is 0.25. Then the mask ratio corresponding to the first iteration is 0.75, the mask ratio corresponding to the second iteration is 0.5, the mask ratio corresponding to the third iteration is 0.25, and the mask ratio corresponding to the fourth iteration is 0.
[0204] In some embodiments, the mask ratio of each round is randomly generated under the condition that the mask ratio is gradually reduced in the order of i from 1 to T. For example, if the mask ratio randomly determined in the first round of iteration is 83%, then a mask ratio can be randomly determined in the interval (0, 0.83) in the second round of iteration. For example, if the mask ratio randomly determined in the second round of iteration is 65%, then a mask ratio can be randomly determined in the interval (0, 0.65) in the third round of iteration, and so on.
[0205] (ii) replacing some word units in the predicted text that meet the mask ratio with the mask to obtain the updated noisy text.
[0206] In some embodiments, the above step (ii) may include: using a random strategy to determine a number of words in the predicted text to be replaced with masks to obtain the updated noisy text, wherein the proportion of the words replaced with masks in the predicted text conforms to the mask ratio.
[0207] The diffusion model may adopt a random strategy in the predicted text to determine one or more word units to replace with masks to obtain the updated noisy text.
[0208] The words predicted and generated in the current iteration of the predicted text constitute a first word-gram set, and the words replaced in the predicted text constitute a second word-gram set, and the second word-gram set is a subset of the first word-gram set. In other words, the words replaced with masks are the words predicted and generated in the current iteration process, and the words generated in the previous iteration process are not replaced. The proportion of words replaced with masks in the predicted text that meets the mask ratio can be that the ratio of words in the second word-gram set to words in the first word-gram set meets the mask ratio.
[0209] As an example, Figure 7B As shown, the word units predicted and generated in the second round of iteration include the second, and the 4th to 12th word units. The first word unit set includes the second, and the 4th to 12th word units, a total of 10 word units. The mask ratio is 80%. 8 word units to be replaced with masks are randomly determined from the word units generated in this round. The word units replaced with masks are the 5th to 12th word units. The second word unit set includes the 5th to 12th word units, a total of 8 word units. The second word unit and the 4th word unit retain the prediction results of this round of iteration.
[0210] In some embodiments, the i-th round of iterations in T rounds of iterations may further include:
[0211] (3) For each word-element in the predicted text that is generated in the current iteration and has not been replaced by a mask, determine that the word-element is generated in the i-th iteration.
[0212] That is to say, among the word-grams generated by this round of iterative prediction, the word-grams that are not replaced by masks have their prediction results determined and will be used as the word-grams in the target text that is finally output.
[0213] As mentioned above, the word units predicted and generated in this round of iteration constitute the first word unit set, and the word units replaced by masks constitute the second word unit set. Therefore, it is determined that the word units generated in this round of iteration are the difference between the first word unit set and the second word unit set.
[0214] For example, in Fig. 7A In , it is determined that the first word unit and the third word unit are generated in parallel in the first iteration. For another example, in Figure 7B In , it is determined that the second word unit and the fourth word unit are generated in parallel during the second iteration. For another example, in Figure 7C In , it is determined that the 9th word unit and the 12th word unit are generated in parallel during the 6th iteration.
[0215] In some embodiments, the target text is divided into multiple text segments, each of which includes a number of word units, and the order of generating word units in the multiple text segments satisfies: if the first word unit in the jth text segment exists in the tth text segment, jIn the j+1th text segment, any second word in the tth text segment is generated in the tth j Round iteration or t j The j is an integer greater than or equal to 1, and the t j is an integer greater than or equal to 1 and less than or equal to the T.
[0216] That is to say, the word in the later text segment is generated later than the word in the previous text segment, or is generated in the same iteration as the word in the previous text segment. Since the word at the front position has the strongest semantic connection with the target prompt text, its reference information is more accurate during the generation process. Therefore, through the above generation method, the overall generation direction of the target text is guaranteed to be from front to back, so that the text generated subsequently can be generated based on the more accurate text generated first, so as to ensure that the text quality is relatively stable during the text generation process.
[0217] Figure 8 A diagram showing a generation process of multiple word-grams in a target text provided according to an embodiment of the present specification is shown. Figure 8 TX in represents a word unit, and M represents a mask. TX can be T1 to T6. For the convenience of illustration, T1 to T6 are shown in different grayscales.
[0218] like Figure 8 As shown, the target text can be divided into three text segments B1, B2 and B3. Each text segment includes four word units. Figure 8 The M in represents the mask, and TX represents the word element determined to be generated in the Xth iteration. Figure 8 Six iterations are shown, and TX may be T1 to T6.
[0219] Among them, the diffusion model generates the first and third words in B1 in parallel in the first round of iteration, generates the second and fourth words in B1 in parallel in the second round of iteration, generates the first and third words in B2 in parallel in the third round of iteration, generates the second and fourth words in B2 in parallel in the fourth round of iteration, generates the first and fourth words in B3 in parallel in the fifth round of iteration, and generates the second and third words in B3 in parallel in the sixth round of iteration.
[0220] In summary, the text generation method provided in this specification uses a trained diffusion model to perform text generation. When generating text, the diffusion model can perform multiple rounds of predictions on the text containing masks of any proportion based on the target prompt text, thereby generating the target text. In each round of prediction, the diffusion model can generate word units in the target text in parallel, thereby improving the efficiency of text generation. Moreover, when generating any word unit in the target text, the diffusion model can use the forward text information and backward text information of the word unit as reference at the same time, so that when generating text, it can perceive the context related to the text and ensure the quality of text generation. Therefore, the text generation method provided in this specification can take into account the quality of text generation while improving the efficiency of text generation.
[0221] On the other hand, this specification provides a computer-readable non-transitory storage medium, storing at least one set of executable instructions for model training or text generation. When the executable instructions are executed by the processor, the executable instructions guide the processor to implement the steps of method P300 or P600 described in this specification. In some possible implementations, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product is run on the system 200, the program code is used to make the system 200 perform the steps of method P300 or P600 described in this specification. The program product for implementing the above method can use a portable compact disk read-only memory (CD-ROM) to include program code and can be run on the system 200. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium containing or storing a program, which can be used by the instruction execution system or used in combination with it. The program product can use any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples of readable storage media include: an electrical connection with one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof. Program code for performing the operations of the present specification may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the system 200, partially on the system 200, as a stand-alone software package, partially on the system 200 and partially on a remote computing device, or entirely on a remote computing device.
[0222] The above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require a specific order or a continuous order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0223] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented only by way of example and may not be limiting. Although not explicitly stated herein, those skilled in the art will appreciate that this specification requires various reasonable changes, improvements and modifications to the embodiments. These changes, improvements and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0224] In addition, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment" and / or "some embodiments" mean that a particular feature, structure or characteristic described in conjunction with the embodiment may be included in at least one embodiment of this specification. Therefore, it can be emphasized and should be understood that two or more references to "an embodiment" or "one embodiment" or "an alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics may be appropriately combined in one or more embodiments of this specification.
[0225] It should be understood that in the foregoing description of the embodiments of this specification, in order to help understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure or its description. However, this does not mean that the combination of these features is necessary. When reading this specification, it is entirely possible for a person skilled in the art to mark out some of the devices as separate embodiments. In other words, the embodiments in this specification can also be understood as the integration of multiple secondary embodiments. And the content of each secondary embodiment is also valid when it is less than all the features of a single aforementioned disclosed embodiment.
[0226] Each patent, patent application, publication of patent applications, and other materials, such as articles, books, specifications, publications, documents, articles, etc., cited herein, except to the extent that it is inconsistent or conflicting with this document or that has a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of a term in any material and the description, definition, and / or use of a term in this document, the term in this document shall prevail.
[0227] Finally, it should be understood that the embodiments of the application disclosed herein are explanations of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are only used as examples and not as limitations. Those skilled in the art can adopt alternative configurations according to the embodiments in this specification to implement the applications in this specification. Therefore, the embodiments of this specification are not limited to the embodiments accurately described in the application.
Claims
1. A text generation method, comprising: Obtaining a target prompt text, where the target prompt text represents the requirements for text generation; And Inputting the target prompt text into a trained diffusion model to generate a target text through the diffusion model, where the target text includes multiple tokens, and the diffusion model is trained to have the ability to perform multiple rounds of prediction on a noisy text based on the target prompt text to gradually generate each token in the target text, where In each round of prediction, the diffusion model generates part of the tokens in the target text in a parallel manner, For a target token in the target text, when generating the target token, the diffusion model uses the forward text information and backward text information of the target token as a reference, where the forward text information includes the text information located before the target token, and the backward text information includes the text information located after the target token.
2. The method according to claim 1, wherein: The generation process of the target text by the diffusion model includes: Generating an initial noisy text, where the initial noisy text includes multiple masks; and Performing T rounds of iteration based on the initial noisy text, where T is an integer greater than 1, and the process of the i-th round of iteration includes: Under the guidance of the target prompt text, performing token prediction on each mask in the current noisy text to obtain a prediction text, where when i = 1, the current noisy text is the initial noisy text, and when i > 1, the current noisy text is the noisy text updated in the (i - 1)-th round of iteration, In the case of i = T, determining the target text based on the prediction text, In the case of i < T, replacing part of the tokens in the prediction text with masks to obtain an updated noisy text.
3. The method according to claim 2, wherein: The replacing part of the tokens in the prediction text with masks to obtain an updated noisy text includes: Determining the mask ratio corresponding to the i-th round of iteration, where, in the order of i from 1 to T, the mask ratio gradually decreases; and Replacing part of the tokens in the prediction text that meet the mask ratio with masks to obtain the updated noisy text.
4. The method according to claim 3, wherein: The tokens predicted and generated in the current round of iteration in the prediction text form a first token set, and the tokens replaced with masks in the prediction text form a second token set, and the second token set is a subset of the first token set.
5. The method according to claim 3, wherein: The determining the mask ratio corresponding to the i-th round of iteration includes: Based on the value of T, determining a target step size; and Performing one of the following based on the target step size to obtain the mask ratio corresponding to the i-th round of iteration: In the case of i = 1, subtracting the target step size from 1, or In the case of i > 1, subtracting the target step size from the mask ratio corresponding to the (i - 1)-th round of iteration.
6. The method according to claim 5, wherein: The method further includes: Obtaining the value of T from the configuration information of the user for the diffusion model.
7. The method according to claim 3, wherein: The replacing part of the tokens in the prediction text that meet the mask ratio with masks to obtain an updated noisy text includes: A random strategy is adopted to determine a number of words in the predicted text to be replaced with masks to obtain the updated noisy text, wherein the proportion of the words replaced with masks in the predicted text conforms to the mask ratio.
8. The method according to claim 2, wherein: The process of the i-th iteration also includes: for each word element in the predicted text that is generated in the current iteration and is not replaced by a mask, determining that the word element is generated in the i-th iteration; The target text is divided into a plurality of text segments, each of which includes a plurality of word units, and the generation order of the word units in the plurality of text segments satisfies: If the first word in the jth text segment exists in the tth j In the j+1th text segment, any second word in the tth text segment is generated in the tth j Round iteration or t j The iteration after the round iteration is generated, wherein j is an integer greater than or equal to 1, and t j is an integer greater than or equal to 1 and less than or equal to the T.
9. The method according to claim 2, wherein: The generating of the initial noisy text comprises: Generate noisy text composed of multiple masks; and Splicing the noise text behind the target prompt text to obtain the initial noisy text; When word-unit prediction is performed on each mask in the current noisy text, at least part of the content in the target prompt text is used as forward text information of each mask.
10. The method according to claim 1, wherein: The diffusion model adopts the Transformer Encoder network architecture and uses a bidirectional attention mechanism in each round of prediction.
11. A model training method, comprising: Obtaining a sample data set, wherein each training sample in the sample data set includes at least one text; Obtain a diffusion model to be trained; as well as The diffusion model is trained for multiple rounds of iterations using the sample data set, wherein each round of iterations includes: Obtain target sample text from the sample data set, A mask ratio is randomly generated in the range of [0,1], and the word units in the target sample text that meet the mask ratio are replaced with masks to obtain noisy text. The diffusion model is used to perform word-unit prediction on each mask in the noisy text to obtain a predicted text, wherein when performing word-unit prediction, the diffusion model predicts the masks in the noisy text in a parallel manner, and for any target mask in the noisy text, when the diffusion model performs word-unit prediction on the target mask, the forward text information and backward text information of the target mask are used as references, the forward text information includes text information located before the target mask, and the backward text information includes text information located after the target mask, The parameters of the diffusion model are updated with the training objective of minimizing the difference between the predicted text and the sample text.
12. The method according to claim 11, wherein: The sample data set includes a first data subset and a second data subset, and the using the sample data set to perform multiple rounds of iterative training on the diffusion model includes: Pre-training the diffusion model using the first data subset so that the diffusion model has the ability to recover clean text from noisy texts with different noise ratios; and The diffusion model is fine-tuned and trained using the second data subset, so that the diffusion model has the ability to generate text under the guidance of the prompt text.
13. The method according to claim 12, wherein: Each training sample in the first data subset includes a text; In the pre-training process, obtaining a target sample text from the sample data set includes: A target training sample is determined from the first data subset, and the text in the target training sample is used as the target sample text.
14. The method according to claim 12, wherein: Each training sample in the second data subset includes a first text and a second text, wherein the first text is used to describe the requirement for text generation, and the second text is a text that meets the requirement; In the process of the fine-tuning training, obtaining the target sample text from the sample data set includes: Determine a target training sample from the second data subset, concatenate the first text and the second text in the target training sample to obtain the target sample text, Wherein, when word-unit prediction is performed on the target mask, at least part of the content in the first text is used as forward text information of the target mask.
15. The method according to claim 14, wherein: The word-grams replaced by masks in the target sample text are word-grams from the second text.
16. The method according to claim 11, wherein: The diffusion model adopts the Transformer Encoder network architecture and uses a bidirectional attention mechanism in the process of word prediction.
17. A text generation system, comprising: at least one storage medium storing at least one instruction set; as well as At least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set when running, and implements the method according to any one of claims 1-10 according to the instructions of the at least one instruction set.
18. A model training system, comprising: at least one storage medium storing at least one instruction set; as well as At least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set when running, and implements the method as described in any one of claims 11-16 according to the instructions of the at least one instruction set.
Citation Information
Cited By
Iterative text refining method, system and equipment based on fidelity and medium
CN120706380A