Training Method for Text Generation Model, Target Corpus Expansion Method and Related Devices
By combining the penalty items of the statistical language model during the training of the text generation model, the problem of failure to effectively utilize the corpus in the existing technology is solved, and the performance of the model is improved.
Patent Information
- Application Number
- CN202111670508.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The prior art fails to effectively utilize the existing corpus when training text generation models, resulting in insufficient model performance.
By obtaining sample corpus, performing word segmentation processing and generating statistical language models, using the generator and discriminator of the text generation model to generate and discriminate the target text, compute the adversarial loss function and confusion, and combine this information to train the text generation model.
By introducing penalties for statistical language models during the training of text generation models, the performance of the model is improved and it can better utilize the existing corpus.
Smart Images

Figure CN114462570B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of text processing, and particularly to a training method for a text generation model, a method for expanding a target corpus, and related devices. Background Art
[0002] Speech recognition includes two parts: an acoustic model and a language model. Among them, the statistical language model is the most widely used in the language model. However, the statistical language model requires a large-scale corpus for support, and in fact, it is difficult to obtain a corpus, especially the corpus in specific scenarios is very scarce. Text automatic generation is an important technology in the fields of artificial intelligence and natural language processing, and it is an effective method to solve the shortage of the corpus. The application scenarios of this technology are extensive, such as common "robot writing", "automatic dialogue generation", "automatic Chinese lyric generator", etc.
[0003] The text generation technology based on deep learning is the mainstream direction in this field. A series of general models have been developed in the development of text generation technology. The existing solution is to first train a text generation model, and use the text generation model to iteratively generate the next word according to the given starting word, without making good use of the existing corpus to guide the training of the text generation model. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a training method for a text generation model, a method for expanding a target corpus, and related devices, which can use the existing corpus to guide the training of the text generation model and improve the performance of the text generation model.
[0005] To solve the above problems, a first aspect of this application provides a training method for a text generation model. The training method for the text generation model includes: obtaining a sample corpus; performing word segmentation processing on the sample corpus, and generating a statistical language model according to the word segmentation processing result; using a generator of the text generation model to generate a target text; using a discriminator of the text generation model to discriminate the target text according to the sample corpus, outputting a discrimination result, and obtaining an adversarial loss function according to the discrimination result; using the statistical language model to obtain the perplexity of the target text, and determining a penalty term according to the perplexity; superimposing the adversarial loss function and the penalty term to obtain a target loss function of the text generation model, and using the target loss function to train the text generation model to obtain a trained text generation model.
[0006] Among them, the tokenization of the sample corpus and the generation of a statistical language model based on the tokenization result include: performing a tokenization operation on the sample corpus using a first tokenization method to obtain a first tokenization result; performing a tokenization operation on the sample corpus using a second tokenization method to obtain a second tokenization result; taking the union of the first tokenization result and the second tokenization result as the tokenization result, and counting the word frequencies in the tokenization result to generate the statistical language model.
[0007] Among them, the first tokenization method includes a dictionary-based method; the performing a tokenization operation on the sample corpus using the first tokenization method to obtain a first tokenization result includes: using the dictionary-based method and adopting a forward maximum matching algorithm to perform a tokenization operation on the sample corpus to obtain a first tokenization result.
[0008] Among them, the second tokenization method includes a method based on a long short-term memory network; the performing a tokenization operation on the sample corpus using the second tokenization method to obtain a second tokenization result includes: performing word embedding processing on the sample corpus, mapping each word in each sentence of the sample corpus into a vector, training word vectors so that the dimension of each word vector is a fixed dimension to process each sentence into a two-dimensional matrix; counting and sorting the word frequencies in the sample corpus according to the two-dimensional matrix, generating an index for each word; dividing each sentence in the sample corpus into word pairs of length n-gram, and converting the word pairs into word indices according to the index generated for each word; training the long short-term memory network using the sample corpus after conversion into word indices to obtain a trained long short-term memory network; performing a tokenization operation on the sample corpus using the trained long short-term memory network to obtain a second tokenization result.
[0009] Among them, the generating of the target text using the generator of the text generation model includes: counting the frequencies of the first words in each sentence of the sample corpus, and taking the top N words with the highest word frequencies as the starting list; selecting a starting word according to the starting list, and calculating all subsequent words through the generator until an end symbol is encountered, generating a new sentence as the target text.
[0010] Among them, the discriminator of the text generation model is used to discriminate the target text according to the sample corpus, output a discrimination result, and obtain an adversarial loss function according to the discrimination result, including: mapping each word in each sentence in the sample corpus into a fixed-dimensional vector through encoding and embedding to represent each sentence as a two-dimensional matrix; wherein, each column in the two-dimensional matrix is composed of the word vectors of the words at the corresponding positions in the sentence; performing convolution and pooling operations on the target text through the discriminator to extract the sentence features of the target text; discriminating the sentence features of the target text according to the two-dimensional matrix corresponding to the sample corpus, outputting a discrimination result of the probability that the target text belongs to the real corpus, and obtaining the adversarial loss function according to the discrimination result.
[0011] Among them, the method for obtaining the perplexity of the target text by using the statistical language model and determining the penalty term according to the perplexity includes: calculating the perplexity of the target text generated by the generator by using the statistical language model, and multiplying the perplexity by a penalty coefficient as the penalty term.
[0012] To solve the above problems, a second aspect of the present application provides a method for expanding a target corpus. The method for expanding the target corpus includes: cleaning the target corpus to obtain a preprocessed target corpus; using a text generation model to expand the preprocessed target corpus; wherein, the text generation model is trained by the training method of the text generation model in the first aspect above.
[0013] To solve the above problems, a third aspect of the present application provides a training device for a text generation model. The training device for the text generation model includes: an acquisition module for acquiring a sample corpus; a word segmentation module for performing word segmentation on the sample corpus and generating a statistical language model according to the word segmentation result; a generation module for generating a target text by using a generator of the text generation model; a discrimination module for discriminating the target text according to the sample corpus by using a discriminator of the text generation model, outputting a discrimination result, and obtaining an adversarial loss function according to the discrimination result; a determination module for obtaining the perplexity of the target text by using the statistical language model and determining a penalty term according to the perplexity; a training module for superimposing the adversarial loss function and the penalty term to obtain a target loss function of the text generation model, and training the text generation model by using the target loss function to obtain a trained text generation model.
[0014] To solve the above problems, a fourth aspect of the present application provides an electronic device. The electronic device for positioning the sound source direction includes a processor and a memory connected to each other. The memory is used to store program instructions, and the processor is used to execute the program instructions to implement the training method of the text generation model in the first aspect above, or the target corpus expansion method in the second aspect above.
[0015] To solve the above problems, a fifth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the training method of the text generation model in the first aspect above, or the target corpus expansion method in the second aspect above is implemented.
[0016] The beneficial effect of the present invention is as follows: Different from the prior art, the present application obtains a sample corpus, then performs word segmentation processing on the sample corpus, and generates a statistical language model according to the word segmentation processing result. Then, a target text is generated by the generator of the text generation model, and then the discriminator of the text generation model is used to discriminate the target text according to the sample corpus, and a discrimination result is output. An adversarial loss function is obtained according to the discrimination result. At the same time, the perplexity of the target text can be obtained by using the statistical language model, and a penalty term is determined according to the perplexity. Then, the adversarial loss function and the penalty term are superimposed to obtain the target loss function of the text generation model, and the text generation model is trained by using the target loss function to obtain the trained text generation model. By obtaining an additional penalty term according to the perplexity of the target text generated by the generator during the training process of the text generation model, and superimposing this penalty term with the original adversarial loss function of the text generation model as the total target loss function to guide the training of the text generation model through backpropagation, the statistical language model can guide the training process of the text generation model, and the performance of the trained text generation model can be improved. Description of the Drawings
[0017] Figure 1 is a schematic flowchart of an embodiment of the training method of the text generation model of the present application;
[0018] Figure 2 is Figure 1 a schematic flowchart of an embodiment of step S12 in
[0019] Figure 3 is Figure 2 a schematic flowchart of an embodiment of step S122 in
[0020] Figure 4 is Figure 1 a schematic flowchart of an embodiment of step S14 in
[0021] Figure 5 is a schematic flowchart of an embodiment of the target corpus expansion method of the present application;
[0022] Figure 6 is a schematic structural diagram of an embodiment of a training device for the text generation model of the present application;
[0023] Figure 7 is a schematic structural diagram of an embodiment of an electronic device of the present application;
[0024] Figure 8 is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. Detailed implementation manners
[0025] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0026] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0027] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article merely describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0028] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of a training method for the text generation model of the present application. The training method for the text generation model in this embodiment includes the following steps:
[0029] Step S11: Obtain sample corpus.
[0030] In the present application, the corpus in a mature corpus publicly available on the network can be used as the sample corpus. The sample corpus can include at least one sentence text, and the at least one sentence text is a real text. In one embodiment, the text in the sample corpus is Chinese text. Therefore, after obtaining the sample corpus, it is necessary to preprocess the sample corpus, mainly to clean the text of the sample corpus, and remove characters such as English, symbols, and numbers, so as to facilitate the subsequent training of the text generation model.
[0031] Step S12: Perform word segmentation on the sample corpus, and generate a statistical language model according to the word segmentation result.
[0032] It can be understood that after performing word segmentation on the sample corpus, a segmented word library can be obtained, and a statistical language model can be made according to the segmented word library.
[0033] Specifically, please combine Figure 2 , Figure 2 is Figure 1 a schematic flowchart of an embodiment of step S12 in
[0034] Step S121: Perform a word segmentation operation on the sample corpus using a first word segmentation method to obtain a first word segmentation result.
[0035] Step S122: Perform a word segmentation operation on the sample corpus using a second word segmentation method to obtain a second word segmentation result.
[0036] Step S123: Take the union of the first word segmentation result and the second word segmentation result as the word segmentation processing result, and count the word frequencies in the word segmentation processing result to generate the statistical language model.
[0037] Specifically, by using the first word segmentation method and the second word segmentation method to perform word segmentation operations on the sample corpus respectively, a first word segmentation result and a second word segmentation result can be obtained correspondingly. Then, take the union operation on the word segmentation results after the two word segmentation methods are completed to obtain the final word segmentation processing result. By counting the word frequencies in the final word segmentation processing result, a statistical language model can be generated, and this statistical language model can be an n-gram language model. By using two word segmentation methods to perform word segmentation processing on the sample corpus and taking the union of the segmented results, the impact on the performance of the obtained statistical language model caused by inaccurate word segmentation in the case of a small-scale sample corpus can be reduced.
[0038] In one embodiment, the above first word segmentation method may include a dictionary-based method, and the above step S121 may specifically include: using the dictionary-based method and adopting a forward maximum matching algorithm to perform a word segmentation operation on the sample corpus to obtain a first word segmentation result.
[0039] Specifically, by using the dictionary-based method, match the sentence to be segmented in the sample corpus with the entries in a dictionary according to a certain strategy. If a certain string is found in the dictionary, it means that the corresponding entry is successfully matched and word segmentation is performed. And when using the forward maximum matching algorithm to perform a word segmentation operation on the sample corpus, specifically, scan from left to right to find the maximum match of the word. First, take the entire sentence to be segmented as the maximum length of a word. Each time when scanning, look for the word with this maximum length starting from the current position to match with the words in the dictionary. If not found, shorten the length and continue to search until found or the length is shortened to a single character.
[0040] Please combine Figure 3 , Figure 3 is Figure 2Schematic flowchart of an embodiment of step S122. In one embodiment, the second word segmentation method includes a method based on a long short-term memory network (LSTM, Long Short-Term Memory). Specifically, step S122 includes:
[0041] Step S1221: Perform word embedding processing on the sample corpus, map each word in each sentence of the sample corpus into a vector, and train the word vectors so that the dimension of each word vector is a fixed dimension, so as to process each sentence into a two-dimensional matrix.
[0042] Step S1222: Count the word frequencies in the sample corpus according to the two-dimensional matrix and sort them, and generate an index for each word.
[0043] Step S1223: Divide each sentence in the sample corpus into word pairs of length n-gram, and convert the word pairs into word indices according to the indices generated for each word.
[0044] Step S1224: Use the sample corpus after being converted into word indices to train the long short-term memory network to obtain a trained long short-term memory network.
[0045] Step S1225: Use the trained long short-term memory network to perform word segmentation on the sample corpus to obtain a second word segmentation result.
[0046] It can be understood that before training the long short-term memory network, it is necessary to preprocess the sample corpus first; specifically, perform word embedding processing on the sample corpus, map each word in each sentence into a vector, and then train the word vectors so that the dimension of each word vector is a fixed dimension. After saving the results, each sentence can be processed into a two-dimensional matrix, and each word in the sentence is an element in the corresponding two-dimensional matrix; then count the word frequencies in the sample corpus according to the two-dimensional matrix and sort them, and generate an index for each word; then, according to the n-gram value selected in the statistical language model, divide each sentence in the sample corpus into word pairs of length n-gram, and convert the word pairs into word indices according to the indices generated for each word to complete the preprocessing of the sample corpus. After that, the sample corpus can be divided into a training set and a test set, use the training set to train the initial long short-term memory network, and use the test set to test the trained long short-term memory network to make the trained long short-term memory network stable; then the trained long short-term memory network can be used to perform word segmentation on the sample corpus to obtain the corresponding second word segmentation result.
[0047] Step S13: Use the generator of the text generation model to generate the target text.
[0048] Step S14: Discriminate the target text using the discriminator of the text generation model based on the sample corpus, output a discrimination result, and obtain an adversarial loss function according to the discrimination result.
[0049] In the embodiment of the present application, the text generation model is a text generation model adopting a generative adversarial network. The text generation model needs to train a generator and a discriminator at the same time. The generator inputs a noise variable and outputs a target text. The discriminator inputs the sample corpus and the target text output by the generator, and outputs a binary classification confidence indicating whether the input is real corpus or forged corpus, that is, it can output the discrimination result of the target text. Ideally, the discriminator needs to accurately judge whether the input data is real corpus or forged corpus, while the generator needs to deceive the discriminator as much as possible so that the discriminator judges all the target texts generated by itself as real corpus. Therefore, during the training process, the discrimination result of the discriminator for the target text may be biased. Thus, the loss value of the adversarial loss function of the text generation model itself can be obtained according to the discrimination result.
[0050] In one embodiment, the above step S13 may specifically include: counting the frequency of the first word of each sentence in the sample corpus, and using the top N words with the highest word frequencies as the starting list; selecting a starting word according to the starting list, and calculating all subsequent words through the generator until an end symbol is encountered, and generating a new sentence as the target text.
[0051] The generator of the text generation model can select a long short-term memory network, which can accurately handle the dependence problem of long sequences. After receiving the noise and passing through the mapping of the network, the long short-term memory network can output a sentence as the target text. Specifically, after inputting the sample corpus into the generator, the generator can count the frequency of the first word of each sentence in the sample corpus, and use the top N words with the highest word frequencies as the starting list. After selecting the starting word according to the starting list, all subsequent words can be calculated using the long short-term memory network in sequence until an end symbol is encountered, and then a new sentence is generated as the target text.
[0052] Please combine Figure 4 , Figure 4 is Figure 1 The flowchart of an embodiment of step S14 in. In one embodiment, the above step S14 specifically includes:
[0053] Step S141: Map each word in each sentence of the sample corpus into a fixed-dimensional vector through encoding and embedding, so as to represent each sentence as a two-dimensional matrix; wherein, each column in the two-dimensional matrix is composed of the word vectors of the words at the corresponding positions in the sentence.
[0054] Step S142: Perform convolution and pooling operations on the target text through the discriminator to extract the sentence features of the target text.
[0055] Step S143: According to the two-dimensional matrix corresponding to the sample corpus, discriminate the sentence features of the target text, output the discrimination result of the probability that the target text belongs to the real corpus, and obtain the adversarial loss function according to the discrimination result.
[0056] It can be understood that before training the text generation model, it is necessary to preprocess the sample corpus, mainly encoding and word embedding work. When the sample corpus is in Chinese, since the model network cannot directly process Chinese, it is necessary to first perform encoding processing on the sample corpus and improve the sparsity after encoding through word embedding. The corresponding word embedding matrix can be gradually learned during the training process of the text generation model. Specifically, the discriminator of the text generation model can select a convolutional neural network. Map each word in each sentence of the sample corpus into a fixed-dimensional vector through encoding and embedding, so that each sentence can be represented as a two-dimensional matrix, and each column in the matrix is composed of the word vectors of the words at the corresponding positions in the sentence. After preprocessing the sample corpus, during the training process of the text generation model, the discriminator performs convolution and pooling operations on the input target text to extract the sentence features of the target text. Finally, the discriminator can output the discrimination result of the probability that the target text belongs to the real corpus, so that the loss value of the adversarial loss function of the text generation model itself can be obtained according to the discrimination result.
[0057] Step S15: Use the statistical language model to obtain the perplexity of the target text, and determine the penalty term according to the perplexity.
[0058] In one embodiment, the above-mentioned Step S15 specifically includes: using the statistical language model to calculate the perplexity of the target text generated by the generator, and multiplying the perplexity by a penalty coefficient as the penalty term.
[0059] Specifically, every time the generator generates a target text, the target text can be restored to the corresponding text through the mapping between word vectors and words. The statistical language model obtained in step S12 can be used to calculate the perplexity of the target text generated by the generator. The perplexity can be used to measure the quality of the target text generated by the generator compared with the real sample corpus. Therefore, after multiplying the perplexity by a penalty coefficient and adding it as a penalty term to the training process of the text generation model, the statistical language model can be used to guide the training of the text generation model.
[0060] Step S16: Superimpose the adversarial loss function and the penalty term to obtain the target loss function of the text generation model, and use the target loss function to train the text generation model to obtain the trained text generation model.
[0061] It can be understood that after superimposing the penalty term and the original adversarial loss function, the target loss function of the text generation model is obtained. Then, using the superimposed target loss function as a whole, through backpropagation, the generator and discriminator of the text generation model are iteratively trained. When the text generation model converges, the iterative process terminates, and the trained text generation model is obtained. The text generation effect of the trained text generation model is better than that of the original text generation model.
[0062] In the above solution, during the training process of the text generation model, an additional penalty term is obtained according to the perplexity of the target text generated by the generator. After superimposing this penalty term and the original adversarial loss function of the text generation model, the total target loss function is used to guide the training of the text generation model through backpropagation, enabling the statistical language model to guide the training process of the text generation model and improving the performance of the trained text generation model.
[0063] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of an embodiment of the target corpus expansion method of the present application. The target corpus expansion method in this embodiment includes the following steps:
[0064] Step S51: Clean the target corpus to obtain the preprocessed target corpus.
[0065] Step S52: Use the text generation model to expand the preprocessed target corpus. Among them, the text generation model is trained by any of the above text generation model training methods.
[0066] In one embodiment, it is necessary to expand the target corpus, and the target corpus is in Chinese. Therefore, after obtaining the target corpus, it is necessary to preprocess the target corpus, mainly to clean the text of the target corpus, removing characters such as English, symbols, and numbers, so as to facilitate subsequent use of the text generation model to expand the target corpus.
[0067] In the training process of the text generation model of the present application, an additional penalty term is obtained according to the perplexity of the target text generated by the generator, and the penalty term is superimposed on the original adversarial loss function of the text generation model as the total target loss function to guide the training of the text generation model through backpropagation, so that the statistical language model can guide the training process of the text generation model, and the performance of the trained text generation model can be improved.
[0068] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of an embodiment of the training device of the text generation model of the present application. The training device 60 of the text generation model in this embodiment includes an acquisition module 600, a word segmentation module 601, a generation module 602, a discrimination module 603, a determination module 604, and a training module 605 that are interconnected; the acquisition module 600 is used to acquire a sample corpus; the word segmentation module 601 is used to perform word segmentation on the sample corpus and generate a statistical language model according to the word segmentation result; the generation module 602 is used to generate a target text using the generator of the text generation model; the discrimination module 603 is used to discriminate the target text using the discriminator of the text generation model according to the sample corpus, output a discrimination result, and obtain an adversarial loss function according to the discrimination result; the determination module 604 is used to obtain the perplexity of the target text using the statistical language model and determine a penalty term according to the perplexity; the training module 605 is used to superimpose the adversarial loss function and the penalty term to obtain the target loss function of the text generation model, and use the target loss function to train the text generation model to obtain the trained text generation model.
[0069] In one embodiment, the step of the word segmentation module 601 performing word segmentation on the sample corpus and generating a statistical language model according to the word segmentation result includes: performing word segmentation on the sample corpus using a first word segmentation method to obtain a first word segmentation result; performing word segmentation on the sample corpus using a second word segmentation method to obtain a second word segmentation result; taking the union of the first word segmentation result and the second word segmentation result as the word segmentation result, and counting the word frequencies in the word segmentation result to generate the statistical language model.
[0070] In one embodiment, the first word segmentation method includes a dictionary-based method. The word segmentation module 601 performs the step of segmenting the sample corpus using the first word segmentation method to obtain a first word segmentation result, specifically including: using the dictionary-based method and adopting the forward maximum matching algorithm to segment the sample corpus to obtain a first word segmentation result.
[0071] In one embodiment, the second word segmentation method includes a long short-term memory network-based method. The word segmentation module 601 performs the step of segmenting the sample corpus using the second word segmentation method to obtain a second word segmentation result, specifically including: performing word embedding processing on the sample corpus, mapping each word in each sentence of the sample corpus into a vector, training word vectors so that the dimension of each word vector is a fixed dimension to process each sentence into a two-dimensional matrix; counting and sorting the word frequencies in the sample corpus according to the two-dimensional matrix to generate an index for each word; dividing each sentence in the sample corpus into word pairs of length n-gram and converting the word pairs into word indices according to the index generated for each word; training the long short-term memory network using the sample corpus after being converted into word indices to obtain a trained long short-term memory network; and using the trained long short-term memory network to segment the sample corpus to obtain a second word segmentation result.
[0072] In one embodiment, the generation module 602 performs the step of generating a target text using the generator of the text generation model, including: counting the frequencies of the first words in each sentence of the sample corpus and taking the top N words with the highest word frequencies as a starting list; selecting a starting word according to the starting list and calculating all subsequent words through the generator until an end symbol is encountered to generate a new sentence as the target text.
[0073] In one embodiment, the discrimination module 603 performs the step of discriminating the target text using the discriminator of the text generation model according to the sample corpus, outputting a discrimination result, and obtaining an adversarial loss function according to the discrimination result, including: mapping each word in each sentence of the sample corpus into a fixed-dimensional vector through encoding and embedding to represent each sentence as a two-dimensional matrix; where each column in the two-dimensional matrix is composed of the word vectors of the words at the corresponding positions in the sentence; performing convolution and pooling operations on the target text through the discriminator to extract the sentence features of the target text; discriminating the sentence features of the target text according to the two-dimensional matrix corresponding to the sample corpus, outputting a discrimination result of the probability that the target text belongs to the real corpus, and obtaining the adversarial loss function according to the discrimination result.
[0074] In one embodiment, the determining module 604 executes the steps of obtaining the perplexity of the target text by using the statistical language model and determining a penalty term according to the perplexity, including: calculating the perplexity of the target text generated by the generator by using the statistical language model, and using the perplexity multiplied by a penalty coefficient as the penalty term.
[0075] For the specific content of the training device 60 of the text generation model of the present application to implement the training method of the text generation model, please refer to the content in the above-mentioned embodiments of the training method of the text generation model, which will not be elaborated here.
[0076] Please refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of an embodiment of an electronic device of the present application. The electronic device 70 in this embodiment includes a processor 702 and a memory 701 connected to each other; the memory 701 is used to store program instructions, and the processor 702 is used to execute the program instructions stored in the memory 701 to implement the steps of any of the above-mentioned embodiments of the training method of the text generation model or the target corpus expansion method. In a specific implementation scenario, the electronic device 70 may include, but is not limited to: a microcomputer, a server.
[0077] Specifically, the processor 702 is used to control itself and the memory 701 to implement the steps of any of the above-mentioned embodiments of the training method of the text generation model or the target corpus expansion method. The processor 702 may also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 702 may be an integrated circuit chip with signal processing capabilities. The processor 702 may also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. In addition, the processor 702 may be implemented jointly by integrated circuit chips.
[0078] Please refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 80 of the present application stores program instructions 800 thereon, and when the program instructions 800 are executed by a processor, the steps in any of the above-mentioned embodiments of the training method of the text generation model or the target corpus expansion method are implemented.
[0079] Specifically, the computer-readable storage medium 80 may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store program instructions 800. Alternatively, it may also be a server storing the program instructions 800. The server can send the stored program instructions 800 to other devices for running, or it can also run the stored program instructions 800 itself.
[0080] In several embodiments provided in the present application, it should be understood that the disclosed methods, devices, and apparatuses can be implemented in other ways. For example, the device and apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0081] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0082] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0083] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
Claims
1. A training method for a text generation model, characterized in that, the training method for the text generation model includes: Obtaining a sample corpus; Performing word segmentation on the sample corpus, and generating a statistical language model according to the word segmentation result; Generating a target text by using a generator of the text generation model; Discriminating the target text by using a discriminator of the text generation model according to the sample corpus, outputting a discrimination result, and obtaining an adversarial loss function according to the discrimination result; Obtaining the perplexity of the target text by using the statistical language model, and determining a penalty term according to the perplexity; Superimposing the adversarial loss function and the penalty term to obtain the target loss function of the text generation model, and training the text generation model by using the target loss function to obtain a trained text generation model.
2. The training method for the text generation model according to claim 1, characterized in that, the performing word segmentation on the sample corpus, and generating a statistical language model according to the word segmentation result includes: Performing word segmentation on the sample corpus by using a first word segmentation method to obtain a first word segmentation result; Performing word segmentation on the sample corpus by using a second word segmentation method to obtain a second word segmentation result; Taking the union of the first word segmentation result and the second word segmentation result as the word segmentation result, and counting the word frequencies in the word segmentation result to generate the statistical language model.
3. The training method for the text generation model according to claim 2, characterized in that, the first word segmentation method includes a dictionary-based method; the performing word segmentation on the sample corpus by using the first word segmentation method to obtain a first word segmentation result includes: Using the dictionary-based method and adopting a forward maximum matching algorithm to perform word segmentation on the sample corpus to obtain a first word segmentation result.
4. The training method for the text generation model according to claim 2, characterized in that, the second word segmentation method includes a method based on a long short-term memory network; the performing word segmentation on the sample corpus by using the second word segmentation method to obtain a second word segmentation result includes: Performing word embedding processing on the sample corpus, mapping each word in each sentence of the sample corpus into a vector, training word vectors so that the dimension of each word vector is a fixed dimension, and processing each sentence into a two-dimensional matrix; Counting and sorting the word frequencies in the sample corpus according to the two-dimensional matrix, and generating an index for each word; Dividing each sentence in the sample corpus into word pairs of length n-gram, and converting the word pairs into word indices according to the index generated for each word; Training the long short-term memory network by using the sample corpus after being converted into word indices to obtain a trained long short-term memory network; Performing word segmentation on the sample corpus by using the trained long short-term memory network to obtain a second word segmentation result.
5. The training method for the text generation model according to claim 1, characterized in that, the generating a target text by using a generator of the text generation model includes: Statistically analyze the frequency of the first word in each sentence of the sample corpus, and use the top N words with the highest word frequencies as the starting list; Select a starting word according to the starting list, and use the generator to calculate all subsequent words until an end symbol is encountered, and generate a new sentence as the target text.
6. The training method of the text generation model according to claim 1, characterized in that the step of using the discriminator of the text generation model to discriminate the target text according to the sample corpus, output a discrimination result, and obtain an adversarial loss function according to the discrimination result, includes: Map each word in each sentence of the sample corpus into a fixed-dimensional vector through encoding and embedding, so as to represent each sentence as a two-dimensional matrix; wherein, each column in the two-dimensional matrix is composed of the word vectors of the words at the corresponding positions in the sentence; Perform convolution and pooling operations on the target text through the discriminator to extract the sentence features of the target text; According to the two-dimensional matrix corresponding to the sample corpus, discriminate the sentence features of the target text, output the discrimination result of the probability that the target text belongs to the real corpus, and obtain the adversarial loss function according to the discrimination result.
7. The training method of the text generation model according to claim 1, characterized in that the step of using the statistical language model to obtain the perplexity of the target text and determining a penalty term according to the perplexity, includes: Use the statistical language model to calculate the perplexity of the target text generated by the generator, and multiply the perplexity by a penalty coefficient as the penalty term.
8. A method for expanding a target corpus, characterized in that the method for expanding the target corpus includes: Perform text cleaning on the target corpus to obtain a preprocessed target corpus; Use a text generation model to expand the preprocessed target corpus; wherein, the text generation model is trained by the training method of the text generation model according to any one of claims 1 to 7.
9. A training device for a text generation model, characterized in that the training device for the text generation model includes: An acquisition module, which is used to acquire a sample corpus; A word segmentation module, which is used to perform word segmentation on the sample corpus and generate a statistical language model according to the word segmentation result; A generation module, which is used to generate a target text by using the generator of the text generation model; A discrimination module, which is used to discriminate the target text by using the discriminator of the text generation model according to the sample corpus, output a discrimination result, and obtain an adversarial loss function according to the discrimination result; A determination module, which is used to use the statistical language model to obtain the perplexity of the target text and determine a penalty term according to the perplexity; A training module, which is used to superimpose the adversarial loss function and the penalty term to obtain the target loss function of the text generation model, and use the target loss function to train the text generation model to obtain a trained text generation model.
10. An electronic device, characterized in that, the electronic device includes a processor and a memory connected to each other; the memory is used to store program instructions, and the processor is used to execute the program instructions to implement the training method of the text generation model according to any one of claims 1-7, or the target corpus expansion method according to claim 8.
11. A computer-readable storage medium, on which program instructions are stored, characterized in that, when the program instructions are executed by a processor, the training method of the text generation model according to any one of claims 1 to 7, or the target corpus expansion method according to claim 8 is implemented.
Citation Information
Patent Citations
Systems and methods for diverse keyphrase generation with neural unlikelihood training
US20220004712A1