Training Method, Text Processing Method and Device of Language Style Conversion Model
By obtaining training sample data, using labels and discriminant model constraints, the language style conversion model is trained, and the existing model is not effective when generating text in specific language styles is solved, achieving a more natural improvement in honorific expression and translation quality.
Patent Information
- Application Number
- CN201910465744.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-05-30
AI Technical Summary
Existing neural network models have poor results in generating text with a specific language style, especially when complex language styles such as honorifics are required, the translation results are stiff and unnatural.
By obtaining training sample data, including the source text of the original language style and the target text of the target language style, and the target text of the non-target language style, the language style transformation model is trained until the total loss function converges, and the output of the model is biased towards the target language style by using labels and discriminant models, and the model performance is improved by using pre-training and retraining methods.
The honorific rate in the translation results is significantly improved, while maintaining the translation quality, ensuring that the output text meets the expression requirements of the target language style.
Smart Images

Figure CN112016271B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology. Specifically, this application relates to a method for training a language style conversion model, a text processing method, and an apparatus. Background Art
[0002] Neural network model architectures represented by encoding (Encode) and decoding (Decode) have been widely used in tasks such as machine translation, text automatic summarization, and robot automatic question answering. Currently, the mainstream architectures of this type mainly use recurrent neural networks plus attention mechanisms or simply use the form of multi-head self-attention. However, when generating text with a specific language style (such as honorifics), the existing models have poor effects. Summary of the Invention
[0003] The purpose of this application is to provide a solution that can generate text with the required language style. To achieve this purpose, the technical solutions provided by this application are as follows:
[0004] In a first aspect, an embodiment of this application provides a method for training a language style conversion model. The method includes:
[0005] Obtain training sample data, where the training sample data includes first training texts and second training samples. Each first training text includes a source text in the original language style and a target text in the target language style corresponding to the source text. Each second training sample includes a source text in the original language style and a target text in a non-target language style corresponding to the source text;
[0006] Train the language style conversion model based on the training sample data until the total loss function of the style conversion model converges. Among them, the total loss function includes a text processing loss function, and the text processing loss function is used to characterize the difference between the text output by the style conversion model and the corresponding target text.
[0007] In a second aspect, an embodiment of this application provides a text processing method. The method includes:
[0008] Obtain the text to be processed;
[0009] Input the text to be processed into the language style conversion model to obtain a target text with the target language style corresponding to the text to be processed, where the language style conversion model is trained based on the method provided in the first aspect of this application.
[0010] In a third aspect, an embodiment of this application provides a training apparatus for a language style conversion model. The apparatus includes:
[0011] A training sample acquisition module, configured to acquire training sample data, where the training sample data includes first training texts and second training samples. Each first training text includes a source text in the original language style and a target text in the target language style corresponding to the source text, and each second training sample includes a source text in the original language style and a target text in a non-target language style corresponding to the source text;
[0012] A model training module, configured to train a language style conversion model based on the training sample data until the total loss function of the style conversion model converges. The total loss function includes a text processing loss function, and the text processing loss function is used to characterize the difference between the text output by the style conversion model and the corresponding target text.
[0013] In a fourth aspect, an embodiment of the present application provides a text processing device, which includes:
[0014] A to-be-processed text acquisition module, configured to acquire a to-be-processed text;
[0015] A target text acquisition model, configured to input the to-be-processed text into the language style conversion model to obtain a target text with the target language style corresponding to the to-be-processed text, where the language style conversion model is trained based on the method provided in the first aspect of the present application.
[0016] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to call the computer program to execute the method provided in the first aspect or the second aspect of the present application.
[0017] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method provided in the first aspect or the second aspect of the present application is implemented.
[0018] The beneficial effects of the technical solutions provided by the present application will be described in detail in combination with specific embodiments and the accompanying drawings hereinafter, and will not be introduced here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.
[0020] Figure 1 Shows a schematic structural diagram of a translation model in the prior art;
[0021] Figure 2 Shows a schematic flowchart of a model training method provided by an embodiment of the present application;
[0022] Figure 3 The figure shows a schematic diagram of the principle of a training method in an example of the present application;
[0023] Figure 4 The figure shows a schematic diagram of the pre-training principle of a translation model in an example of the present application;
[0024] Figure 5 The figure shows a schematic diagram of the pre-training principle of a discriminant model in an example of the present application
[0025] Figure 6 The figure shows a schematic diagram of the re-training principle of a translation model in an example of the present application;
[0026] Figure 7 The figure shows a schematic diagram of the principle of calculating a first discriminant loss function;
[0027] Figure 8 The figure shows a schematic diagram of the structure of a training device for a model provided by an embodiment of the present application;
[0028] Figure 9 The figure shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0029] Figure 10 The figure shows a schematic diagram of the principle of calculating a loss function provided by an embodiment of the present application. Detailed implementation manners
[0030] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and cannot be construed as a limitation of the present invention.
[0031] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0032] To make the objectives, technical solutions, and advantages of this application clearer, the following takes a translation model as an example to briefly introduce the existing related technical solutions.
[0033] Figure 1 FIG. shows a schematic structural diagram of an existing machine translation model. As shown in the figure, the model mainly includes two parts: an encoder and a decoder. For the text to be translated, such as the Chinese text "I am a student" shown in the figure, input this text into the encoder of the model, and the decoder can output the translation text corresponding to the target language style (Korean in this example), such as the Korean text in the output part shown in the figure.
[0034] Although the existing machine translation models can also relatively well achieve text translation, the applicant of this application has found that for languages with honorific expressions, the existing machine translation models cannot ensure that the translation results are expressed in an honorific way. Specifically, honorifics in languages are very complex. For example, in languages such as Japanese and Korean, nouns and verbs have special honorific forms, and there are also morphological changes at the end of sentences. Simple word substitution (the existing common translation scheme) will make the generated sentences rather rigid. And when the target language is a language with honorific expressions, there are likely to be different honorific expression ways in different context. How to naturally generate honorifics under different task frameworks (translation or Q&A) is still an urgent problem to be solved.
[0035] Similarly, in addition to translation models, other commonly used existing neural network models basically cannot meet the application requirements of generating target texts with specific language styles.
[0036] In addition, it can be understood that language style refers to the specific way of expressing text with a certain meaning. For example, language styles can be honorifics, non - honorifics, colloquial, written, etc.
[0037] To solve the above problems existing in the prior art, this application provides a training method, a text processing method, and a device for a language style conversion model. Based on the solution provided by this application, it is possible to generate target texts with the target language style. The following specifically describes the solution of this application with specific embodiments.
[0038] Figure 2 FIG. shows a schematic flow diagram of a training method for a language style conversion model provided by an embodiment of this application. As shown in the figure, the training method mainly may include the following steps:
[0039] Step S110: Obtain training sample data. The training sample data includes first training texts and second training samples. Each first training text includes a source text in the original language style and a target text in the target language style corresponding to the source text. Each second training sample includes a source text in the original language style and a target text in a non-target language style corresponding to the source text.
[0040] That is to say, each training sample in the training sample data is a pair of texts, namely the source text and the target text corresponding to this text, and the training data includes both target texts in the target language style and target texts in non-target language styles. For example, if the target language style is honorific language, then the target samples include honorific texts and non-honorific texts.
[0041] Among them, the language style conversion model refers to a model with the function of text language style conversion. However, the specific type of the model is not limited in this application. For different application requirements, the model can be a functional model corresponding to the actual application requirements. For example, the style conversion model can include, but is not limited to, a translation model or a question-and-answer model, etc.
[0042] It can be understood that for different such style conversion models, the training sample data can also be different. For example, for a translation model, the language types of the source text and the target text in its training sample data are different. At this time, the source text is the text corresponding to the source language (which can be simply referred to as the source language text), and the target text is the text corresponding to the target language (which can be simply referred to as the target language text). For example, if the translation model to be trained is a model that translates English into Korean, then English is the source language and Korean is the target language, and each training text in the training sample data includes an English text and the Korean text corresponding to this English text; again, for a question-and-answer model, the source text is the question text and the target text is the corresponding answer text.
[0043] It can be seen that in the embodiments of this application, the source text and the target text can be texts in the same language or texts in different languages. In addition, the source text and the target text can also be two texts with the same essential content, that is, the target text is an expression of the source text in another language style. That is to say, the language style conversion model in the embodiments of this application can be a model with language style conversion function and other functions (such as translation, question-and-answer, etc.), or can be just a model with language style conversion function. In addition, for models that have other functions in addition to language style conversion, the original language style and the target language style can be the same or different.
[0044] Step S120: Train the language style conversion model based on the training sample data until the total loss function of the style conversion model converges.
[0045] During the training phase, the input of the language style conversion model is the source text, and the output is the predicted target text. For example, for a translation model, its input is the source language text, and the output is the translation result of the source language text, that is, the predicted target text.
[0046] In the embodiments of the present application, the total loss function includes a text processing loss function, which is used for the difference between the text output by the style conversion model and the corresponding target text (the target text in the training sample). Since the target text in the training sample data includes both the text in the target language style and the text not in the target language style, through model training, the output result of the model can be biased towards the expression in the target language style.
[0047] Among them, the specific form of the text processing loss function is not limited in the embodiments of the present application and can be selected according to requirements. As an optional solution, during each training, for a sample data, the value of its corresponding text processing loss function can be obtained by calculating the difference between the text output by the model and the target text in the sample data.
[0048] In the optional embodiments of the present application, training the language style conversion model based on the training sample data includes:
[0049] Annotating the first label of the source text in the training sample data, where the first label is used to represent whether the language style of the source text is the target language style;
[0050] Training the style conversion model based on the labeled source text and the corresponding target text.
[0051] As an optional solution, during the training phase of the style conversion model, the labels of each source text in the sample data can be annotated by setting a standard or manually. Through this label, it can be known whether the corresponding target text of the source text is a text with the target language style. For example, for a translation model, assuming the target language style is honorific language, this label is used to indicate whether the target language text contains honorific language.
[0052] Among them, the specific form of label annotation can be set according to requirements. As an optional method, a label can be added at the end of the source text. For example, if the target text corresponding to the source text is in the style of the target language, the label "P" can be added at the end of the source text. If the target text corresponding to the source text is in the style of the target language, the label "np" can be added at the end of the source text. Additionally, in actual applications, in order to make the features of the label more prominent when the model processes the text, some special processing can be performed on the text itself and the label. For example, a space or an underscore (such as "_") can be added between the text itself and the label. Taking the translation model and the honorific style of the target language as an example, if the target language text contains honorifics, the label "_p" can be added at the end of the source language text before training. If the target language text does not contain honorifics, the label "_np" can be added at the end of the corresponding source language text.
[0053] Based on the training method with labeled data, after the model is trained, when using the model to process the text to be processed, the output result of the model can be biased towards the result with the style of the target language.
[0054] As an example, Figure 3 FIG. shows a schematic flow chart of a training method for labeled data provided by an embodiment of the present application. In this example, the language style conversion model is a translation model, and the target language style is the honorific style. As shown in the figure, the part indicated by the dashed box in this example is the translation model (in the example of the present application, for the convenience of describing different parts of the model, the word embedding model of the language style conversion model is separately illustrated from other parts of the model, but other parts are still referred to by the name of the entire model. For example, the translation model in this example is actually the part of the translation model except the word embedding part of the input segment). The data_p in the figure represents that the target language text in the sample data contains honorifics, and the data_np represents that the target language text does not contain honorifics. When training the model, the source language text in each training sample is input into the word embedding model, and the translation model encodes and decodes according to the word vectors output by the word embedding model, and outputs the corresponding translation result. During each training process, it is determined whether the training is completed based on whether the total loss function converges.
[0055] Based on the training method with labeled data, after the model is trained, when using the model for translation, for the text to be translated, a label can be annotated on the text to be translated before inputting it into the model. For example, the label "p" can be added at the end of the text to be translated, and the labeled text to be translated is input into the trained model. Based on this method, the translation result output by the model can be biased towards the result containing honorifics.
[0056] In an alternative embodiment of the present application, the language style conversion model is trained based on training sample data until the total loss function of the style conversion model converges, including:
[0057] A language style discriminant model is set up. This discriminant model is used to determine the probability that the text output by the style conversion model is text with the target language style. Among them, the total loss function also includes a first discriminant loss function corresponding to the discriminant model;
[0058] Based on the training sample data and the discriminant model, the style conversion model is trained until the total loss function converges.
[0059] In this solution, by adding a discriminant model during the training process of the style conversion model to constrain the processing result of the conversion model, it can be ensured that when the trained style conversion model is applied, the language style of the output result of the model tends to the target language style, such as tending to an expression containing honorifics.
[0060] Correspondingly, in order to ensure the output effect of the trained style conversion model, when adding a discriminant model, a corresponding first discriminant loss function is also added to improve the discrimination effect of the discriminant model and the processing effect of the style conversion model through this loss function.
[0061] It can be understood that the essence of the discriminant model can be a classification network, and its function is to determine whether the text corresponding to its input is text with the target language style. As an alternative, the discriminant model can be a convolutional neural network (CNN) classifier.
[0062] Among them, the specific function form of the first discriminant loss function can be selected according to actual needs. For example, a discriminant loss function commonly used in neural network training, that is, a classification loss function, can be selected. As an alternative, the first discriminant loss function can be determined according to the discriminant result of the discriminant model and the text corresponding to the input of the discriminant model corresponding to this discriminant result. That is, the value of this function can be determined by the probability that the result discriminated by the discriminant model is text with the target language style and the actual situation of whether the text corresponding to the input of the discriminant model is text with the target language style.
[0063] In an alternative embodiment of the present application, the value of the first discriminant loss function is determined based on the scores of each candidate output of the style conversion model and the probability that the discriminant result corresponding to each candidate output is text with the target language style.
[0064] For example, for a translation model, if the target language style is an honorific style, the value of the first discrimination loss function can be determined based on the translation scores of the candidate translation results of the translation model and the probabilities that the discrimination results corresponding to the candidate translation results are honorifics.
[0065] That is to say, during the actual training process, multiple candidate output results of the style conversion model can be selected, and the discrimination model can discriminate the selected multiple candidate output results respectively. The value of the first discrimination loss function can be calculated from the scores of the multiple candidate output results and the corresponding discrimination results.
[0066] As an optional solution, the total loss function Traing Loss can be expressed as:
[0067]
[0068] Among them, both γ and (1 - γ) are adjustment coefficients, which can be configured and adjusted according to actual needs, 0 ≤ γ ≤ 1, N is the number of samples, that is, the number of source texts input into the model each time during training, Loss(N) can be a common sample processing loss function, and γ * Loss(N) represents the text processing loss function in this optional method. represents the first discrimination loss function corresponding to a sample, K represents the number of candidate output results, score(f g_i ) represents the score of the i-th candidate output result. represents the discrimination result corresponding to the i-th candidate output result, that is, the probability that the text corresponding to the i-th candidate output result is a text with the target language style. represents taking the logarithm operation.
[0069] In an optional embodiment of the present application, training the language style conversion model based on training sample data until the total loss function of the style conversion model converges includes:
[0070] Pre-training the translation model based on training sample data until the text processing loss function converges;
[0071] Training the pre-trained style conversion model based on training sample data until the total loss function converges.
[0072] That is to say, before training the style conversion model, the initial style conversion model can be pre-trained based on the text processing loss function, and then, based on the total loss function, the pre-trained style conversion model can be trained. Through the pre-training step, the text processing accuracy of the model can be improved, that is, the accuracy of the content of the target text output by the model can be improved. On the basis of relatively good accuracy after pre-training, training the model can also reduce the training time and improve the training speed of the model.
[0073] In an optional embodiment of the present application, training the language style conversion model based on the training sample data until the total loss function of the style conversion model converges includes:
[0074] Pre-training the discriminant model based on the output of the pre-trained style conversion model until the second discriminant loss function converges;
[0075] Training the pre-trained style conversion model based on the training sample data and the pre-trained discriminant model until the total loss function converges.
[0076] That is to say, after completing the pre-training of the style conversion model, before training the style conversion model again, the discriminant model can also be pre-trained so that the pre-trained discriminant model can better distinguish whether the text corresponding to its input is in the target language style or non-target language style. Based on this solution, the accuracy of the discriminant model used in subsequent training of the translation model can be effectively improved, the efficiency of training the style conversion model again can be improved, and the text processing effect of the finally trained style conversion model and the probability of outputting text with the target language style can be further improved.
[0077] In an optional embodiment of the present application, the value of the second discriminant loss function is determined based on the discriminant result of the discriminant model and the second label of the output of the style conversion model corresponding to the discriminant result, where the second label is used to represent whether the output of the style conversion model is an output with the target language style.
[0078] As an optional method, the second discriminant loss function can adopt the loss function commonly used in neural network training, that is, the value of the function can be calculated according to the discriminant result of the model and the actual situation of whether its input (that is, the output of the segmentation conversion model) is an input with the target language style.
[0079] As an optional method, the second discriminant loss function Loss(C) can be expressed as:
[0080]
[0081] where M represents the number of inputs of the discriminant model during each training, yj A label representing the input of the discrimination model, which characterizes whether the input has the target language style, y j The value range of y can be {0, 1}. Specifically, y j When the value of y is 0, it means that the input does not have the target language style, that is, it is a non-target language style, y j When the value of y is 1, it means that the input has the target language style, y p Represents the discrimination result of the discrimination model, that is, whether the output of the model-determined style conversion model has the output of the target language style.
[0082] In an alternative embodiment of the present application, the discrimination model is pre-trained based on the output of the pre-trained style conversion model until the second discrimination loss function converges, including:
[0083] Fix the model parameters of the pre-trained style conversion model, input the source text in the training sample data into the pre-trained style conversion model, and train the discrimination model based on the word vectors corresponding to the output of the pre-trained style conversion model until the second discrimination loss function converges.
[0084] That is to say, after the pre-training of the style conversion model is completed, when the discrimination model is pre-trained, since the style conversion model is already a pre-trained model, the model parameters of the style conversion model can be fixed at this time, and the discrimination model can be trained based on the word vectors corresponding to the output text of the model, and the parameters of the discrimination model can be continuously adjusted during the training process until the pre-training of the discrimination model ends when the second discrimination loss function converges.
[0085] It can be understood that in practical applications, in addition to using the above word vectors as the input of the discrimination model to pre-train the discrimination model, the output result of the style conversion model can also be used to train the discrimination model.
[0086] As an alternative, when pre-training the discriminative model after completing the pre-training of the style conversion model, part or all of the source samples in the training samples can be used to train the discriminative model. That is to say, part or all of the source corpora (i.e., the source texts in the samples) used in the training of the style conversion model can be extracted. For each source corpus, it is input into the pre-trained style conversion model. The N1 (the best first N1, i.e., the first best-N1) outputs with high scores of the style conversion model can be selected to train the discriminative model. Specifically, for the first best-N1 outputs of each source corpus output by the style conversion model, second labels are marked. For example, labels 0 and 1 respectively represent that the output is a non-target language segmentation output and a target language style output. Based on these labeled outputs, the discriminative model is trained until the second loss function converges. In this example, M in the above formula is the product of the number of source corpora extracted for training the discriminative model and N1.
[0087] For example, for a translation model, taking the target language style as honorific language as an example, a part can be extracted from the source corpora used for pre-training neural machine translation (NMT). The best-N1 (when a source corpus is input, NMT will have multiple candidate outputs) translation results of the pre-trained NMT for each source corpus are taken, and then the translation results are marked with honorific language. 0 and 1 respectively represent non-honorific language and honorific language, and then they are used to train the honorific language discriminative model. Among them, the value of N1 can be configured according to the actual situation. For example, it can be set to 5.
[0088] In an alternative embodiment of the present application, based on the training sample data and the pre-trained discriminative model, the pre-trained style conversion model is trained, including:
[0089] Fix the model parameters of the pre-trained discriminative model, input the source text in the training sample data into the pre-trained style conversion model, and obtain the output of the pre-trained style conversion model;
[0090] Input the word vectors corresponding to the output of the pre-trained style conversion model into the pre-trained discriminative model to obtain the corresponding discriminative results.
[0091] Since the discriminative model is a pre-trained model, therefore, when the style conversion model is re-trained subsequently, the model parameters of the discriminative model can be fixed, and the re-training of the style conversion model is constrained based on the text processing loss function and the first discriminative loss function, improving the processing effect of the finally obtained style conversion model. On the basis of ensuring the accuracy of text content processing, the possibility of the model outputting text with the target language style is improved.
[0092] In an alternative embodiment of the present application, the style conversion model includes a word embedding model. Fix the model parameters of the discriminant model after pre-training, including:
[0093] The model parameters of the discriminant model after pre-training and the parameters of the word embedding model.
[0094] In practical applications, the word embedding model generally includes a word embedding model at the input end of the model (which can be called Inputsembedding, input embedding) and a word embedding model at the output end (called Outputsembedding, output embedding). For example, Figure 3 The word embedding shown in is the input embedding. Among them, the input embedding is used to map each character or word in the input text (the source text during model training, the text to be processed during model use) into a word vector of a fixed dimension. The output embedding is used to map the character or word predicted by the previous step of the decoder into a word vector of a fixed dimension. The decoder can predict the current character or word based on this word vector and the output of the encoder.
[0095] When pre-training the style conversion model, as part of the style conversion model, the parameters of the word embedding model are already pre-trained parameters. Therefore, when training the translation model again, the parameters of the word embedding model can be fixed without further training, and the parameters of other network structure parts of the style conversion model can be trained again to improve the processing effect of the model. In addition, since the discriminant model is pre-trained based on the pre-trained word embedding model, that is, the word vectors used by the discriminant model during pre-training are obtained based on the pre-trained word embedding model. Therefore, fixing the parameters of the word embedding model during re-training can not only affect the discriminant model but also improve the model training speed.
[0096] To better illustrate and understand the solution provided by the embodiments of the present application, the following is a further description with examples.
[0097] In this example, taking the style conversion model as a translation model and the target language style as honorific language as an example, the solution provided by the embodiments of the present application is described.
[0098] Figure 4 shows a schematic diagram of the translation model in this example. Figure 5 shows a schematic diagram of the discriminant model in this example. Figure 6 shows a schematic diagram of the principle of training the translation model based on the discriminant model in this example. As shown in Figure 4 As shown in, the translation model in this example includes a word embedding model (only the input word embedding is shown in the figure) and a translation model. Figure 5 and Figure 6The honorific discriminator model (Discriminator, denoted as D in the figure) shown in the figure is the discriminator model in this example, which is composed of Figure 6 As can be seen from Figure 6 , the input of the discriminator model is the output of the output word embedding layer of the translation model (the output labeled with honorific tags or non-honorific tags).
[0099] In this example, the translation loss function adopts a general translation loss function, denoted as Loss(N), and the second discriminant loss function is denoted as Loss(C). Formula (2) in the previous text can be used, that is:
[0100]
[0101] The first discriminant loss function can be expressed as in formula (1) in the previous text, that is:
[0102]
[0103] The training process in this example will be described in detail below:
[0104] Step 1: As shown in Figure 4 , pre-train the translation model, that is, pre-train the translation model structure shown in Figure 3 . Figure 4 Specifically, the loss function used in this pre-training step can directly adopt the translation loss function Loss(N). Input the source language samples in the training sample data (the training data shown in the figure) into the translation model, and train the model until Loss(N) converges. Through this pre-training step, a translation model that initially meets the translation quality can be obtained. Figure 3 As shown in , pre-train the translation model structure shown in
[0105] .
[0105] It should be noted that in this pre-training step, the model parameters to be trained include the parameters of the word embedding model of the translation model and the model parameters of other parts (the translation model shown in the figure). The parameters of the word embedding model in the translation model after pre-training are used as the parameters of the word embedding model in the subsequent training process. That is to say, after completing this pre-training step, the parameters of the word embedding model in the subsequent training process are fixed.
[0106] It should be noted that in this pre-training step, the model parameters to be trained include the parameters of the word embedding model of the translation model and the model parameters of other parts (the translation model shown in the figure). The parameters of the word embedding model in the translation model after pre-training are used as the parameters of the word embedding model in the subsequent training process. That is to say, after completing this pre-training step, the parameters of the word embedding model in the subsequent training process are fixed.
[0107] Step 2: As shown in Figure 5 , in this step, pre-train the honorific discriminator model. Through the style conversion model obtained in Step 1, train the discriminator model by fixing the translation model (that is, fixing the model parameters of the word embedding part and other parts in the translation model), so that the discriminator model can correctly distinguish honorifics and non-honorifics in the end. Figure 5 As shown in , in this step, pre-train the honorific discriminator model. Through the style conversion model obtained in Step 1, train the discriminator model by fixing the translation model (that is, fixing the model parameters of the word embedding part and other parts in the translation model), so that the discriminator model can correctly distinguish honorifics and non-honorifics in the end.
[0108] Specifically, when pre-training the discriminant model based on the second discriminant loss function, i.e., Loss(C), the input of the discriminant model is the output of the output embedding layer of the pre-trained translation model. That is to say, Figure 5 the word embedding shown in Figure 5 is the output embedding of the translation model. As shown in
[0109] Step 3: As shown in Figure 6 this step retrains the translation model.
[0110] Specifically, based on the translation model pre-trained in Step 1, and at the same time, the honorific discriminant model uses the discriminator model pre-trained in Step 2. During the retraining process, the parameters of the discriminant model are fixed, and the parameters of the word embedding model of the translation model are fixed (i.e., using the fixed word embedding method). Finally, the combined loss of the discriminant model (corresponding to the first discriminant loss function) and the loss of the translation model (corresponding to the text processing loss function) are used to update the parameters of the translation model. In this step, the total loss function Traing Loss in the above formula (1) can be used for retraining.
[0111] Figure 7 shows a schematic diagram of how to obtain the corresponding character, word, or sub-word (i.e., the token at the target end) according to the word vector after obtaining the word vector corresponding to the translation result through the decoder of the translation model in an example of the present application. In practical applications, for the output of the decoder, it can first pass through a linear layer (Linear, i.e., a fully connected neural network) to map the vector output by the decoder into a longer vector. For example, if the size of the target end dictionary is V, then the dimension of the vector output by the linear layer is V. The vector input to the linear layer passes through the softmax layer shown in the bottom layer of the figure, and the softmax layer converts the vector into a probability vector (i.e., obtaining the output probability through softmax, i.e., OutputProbabilities). The element value of the i-th dimension of this probability vector represents the probability that the predicted character or word is the corresponding character or word in the dictionary for this dimension.
[0112] In this example, it is assumed that the sentence length is T (that is, the text consists of T characters and / or words), and the candidate translation results of the translation model are k, as shown in Figure 7 As shown, the translation model is a process of cyclic decoding, which continuously generates tokens (characters or words in the target text) at the target end (i.e., the decoding end). When decoding for the i-th time, the i-th character or word needs to be generated. After passing through softmax, a probability vector of dimension V is obtained. The elements in this probability vector can be sampled k times (i.e., sampling K times) by multinomial sampling with replacement, resulting in k sampling results, and the corresponding tokens for each result can be obtained (the K tokens shown in the figure). In this way, after T samplings of k times, for each of the T tokens in the target text, k candidate tokens can be obtained. As shown in the figure, for the first token, its corresponding k candidate tokens can be represented as tok1_1, tok1_2,..., tok1_k, and for the T-th token, its corresponding k candidate tokens can be represented as tokT_1, tokT_2,..., tokT_k. After obtaining the candidate tokens, k candidate translation results, that is, k candidate sentences, can be obtained, such as scorel to scorek shown in the figure. It can be understood that for the i-th candidate sentence, this sentence is obtained from tok1_i, tok2i,..., tokT_i.
[0113] After determining the candidate sentences, the translation score of each sentence can be calculated based on the probability values corresponding to the T candidate tokens of each candidate sentence. For example, for the first candidate sentence scorel, its translation is Ptok1_1*Ptok2_1*…*PtokT_1 shown in the figure, where Ptoki_j represents the probability value corresponding to the result of the j-th sampling when generating the i-th token. For example, Ptok1_1 represents the probability value corresponding to the result of the first sampling when generating the first token. That is, multiplying the probabilities corresponding to each token in each candidate sentence can obtain the translation score of this sentence. As an optional method, after obtaining the scores of each candidate sentence, a softmax process can also be used to obtain the relative scores among the candidate sentences, that is, to obtain the normalized scores of each candidate sentence.
[0114] For each candidate sentence, the discriminator is used to judge the probability value that each candidate sentence is an honorific, as shown in the figure indicating the probability value that the i-th candidate sentence contains an honorific. Based on this probability value and the normalized score of each sentence, the first discriminant loss function can be calculated ( Figure 7The value of loss(C) shown in
[0115] It should be noted that, as an optional method, when calculating the first discriminant loss function, for each candidate sentence, the score of the candidate sentence obtained by multiplying the probabilities can be used, or the above-mentioned normalized score can be used. Using the normalized score can make the scores involved in the calculation relatively more objective.
[0116] For those skilled in the art, it is clear that in practical applications, as an optional solution, during the cyclic decoding process, the token corresponding to the maximum element value in the Output Probabilities obtained each time can be obtained, and the T tokens with the highest probabilities can be obtained. The sentences corresponding to these T tokens can be used as the final translation results of the translation model, and Loss(N) can be calculated based on the T tokens and the corresponding target language text in the sample data.
[0117] As shown in Figure 6 During the retraining process, the overall loss function Training Loss determined by the first discriminant loss function and the translation loss function (the overall loss function shown in formula (1)) can be used to guide the further training of the translation model until Training Loss converges, and the translation model obtained from this training is used as the final translation model.
[0118] Compared with the prior art, the solution for training the style conversion model based on the discriminant model provided by the embodiments of the present application introduces a method of minimum risk training based on the discriminant model and maximum likelihood training of the style conversion model during the training process of the style conversion model. By introducing a pre-trained discriminator (which can determine whether the translation result is honorific or non-honorific) and adding it to the normal training of the style conversion model, that is, fusing the loss of the discriminant model into the loss of the style conversion model. During the training process, if the output result of the style conversion model is a result with the target language style (such as the translation result is honorific), the loss of the discriminant model is small, so the overall loss of the model during training is small, and the model parameters are updated little; if the output result is a result with a non-target language style, the loss of the discriminant model is large, so the overall loss of the model during training is large, and the model parameters are updated greatly. The style conversion model obtained based on this training method will tend to have the target language style when using the model in the online stage.
[0119] The solution provided by the embodiments of the present application is applicable to various general application scenarios that require outputting text in a target language style. For example, for a translation model, it can be applicable to various text translation scenarios where the target language is text with a specific language style. For instance, based on this solution, the model can output text in the target language with honorifics, such as Korean, Japanese, etc. The solution of the embodiments of the present application is applicable to the transformation of text styles, and similar solutions can also be adopted for different styles.
[0120] As Figure 10 shown, it is a schematic diagram of calculating the loss function proposed by the embodiments of the present application. Inputting a Token, it is first converted into a corresponding word embedding, and through NMT, the corresponding text processing loss function is obtained and the NMT output is predicted. The NMT output is sampled K times to obtain K sampled sentences, the scores of the K sentences are calculated, and at the same time, the K sentences are input into a discriminant model to predict the honorific probability values of each sentence. Finally, the negative logarithm of the K sentences is calculated and multiplied by the scores of the K sentences to obtain the first discriminant loss function.
[0121] For the solution provided by the embodiments of the present application, taking a translation model that needs to output text in an honorific style as an example, assuming the source language is Chinese and the target language is Korean, 2133 Chinese texts are randomly selected to test the existing translation model and the translation model provided by the embodiments of the present application. The test results are shown in the following table:
[0122] BLEU Honorific rate Existing translation model 34.21 44.77% First solution of this application 33.57 79.23% Second solution of this application 34.17 99.40%
[0123] Among them, Solution 1 refers to the translation model trained based on the sample data with honorific labels and non - honorific labels provided in the embodiments of the present application, and Solution 2 refers to the translation model trained by training the translation model based on the training sample data and the set discriminant model.
[0124] As shown in the table, BLEU (bilingual evaluation understudy) is the translation quality evaluation standard, and the higher the score, the higher the translation quality. The BLEU and honorific rate in the test results shown in this table are calculated as the mean of the test results of the top 5 BLEU scores obtained from the test. It can be seen from the table that based on the solution provided by the embodiments of the present application, the honorific rate in the translation results can be significantly improved while the BLEU remains basically unchanged.
[0125] Based on the same principle as the method Figure 2 shown in, the embodiments of the present application also provide a training device for a language style conversion model. As Figure 8 shown in, the training device 100 of the language style conversion model may include a training sample acquisition module 110 and a model training module 120. Among them:
[0126] The training sample acquisition module 110 is used to acquire training sample data, where the training sample data includes first training texts and second training samples. Each first training text includes a source text in the original language style and a target text in the target language style corresponding to the source text. Each second training sample includes a source text in the original language style and a target text in a non-target language style corresponding to the source text;
[0127] The model training module 120 is used to train the language style conversion model based on the training sample data until the total loss function of the style conversion model converges. Among them, the total loss function includes a text processing loss function, and the text processing loss function is used to characterize the difference between the text output by the style conversion model and the corresponding target text.
[0128] Optionally, the language style conversion model includes a translation model or a question-and-answer model.
[0129] Optionally, the target language style includes an honorific style.
[0130] Optionally, when the model training module 120 trains the language style conversion model based on the training sample data, it can specifically be used for:
[0131] Label the first label of the source text in the training sample data, where the first label is used to characterize whether the language style of the source text is the target language style;
[0132] Train the style conversion model based on the labeled source text and the corresponding target text.
[0133] Optionally, when the model training module 120 trains the language style conversion model based on the training sample data until the total loss function of the style conversion model converges, it is specifically used for:
[0134] Set up a language style discrimination model. The discrimination model is used to determine the probability that the text output by the style conversion model is a text with the target language style. Among them, the total loss function also includes a first discrimination loss function corresponding to the discrimination model;
[0135] Train the style conversion model based on the training sample data and the discrimination model until the total loss function converges.
[0136] Optionally, the value of the first discrimination loss function is determined based on the scores of each candidate output of the style conversion model and the probability that the discrimination result corresponding to each candidate output is a text with the target language style.
[0137] Optionally, when the model training module 120 trains the language style conversion model based on the training sample data until the total loss function of the style conversion model converges, it is specifically used for:
[0138] Pre-train the translation model based on the training sample data until the text processing loss function converges;
[0139] Train the pre-trained style conversion model based on the training sample data until the total loss function converges.
[0140] Optionally, when the model training module 120 trains the language style conversion model based on the training sample data until the total loss function of the style conversion model converges, it is specifically used for:
[0141] Pre-train the discriminant model based on the output of the pre-trained style conversion model until the second discriminant loss function converges;
[0142] Train the pre-trained style conversion model based on the training sample data and the pre-trained discriminant model until the total loss function converges.
[0143] Optionally, the value of the second discriminant loss function is determined based on the discriminant result of the discriminant model and the second label corresponding to the output of the style conversion model corresponding to the discriminant result, where the second label is used to characterize whether the output of the style conversion model is an output with the target language style.
[0144] Optionally, when the model training module 120 pre-trains the discriminant model based on the output of the pre-trained style conversion model until the second discriminant loss function converges, it is specifically used for:
[0145] Fix the model parameters of the pre-trained style conversion model, input the source text in the training sample data into the pre-trained style conversion model, and train the discriminant model based on the word vectors corresponding to the output of the pre-trained style conversion model until the second discriminant loss function converges.
[0146] Optionally, when the model training module 120 trains the pre-trained style conversion model based on the training sample data and the pre-trained discriminant model, it is specifically used for:
[0147] Fix the model parameters of the pre-trained discriminant model, input the source text in the training sample data into the pre-trained style conversion model, and obtain the output of the pre-trained style conversion model;
[0148] Input the word vectors corresponding to the output of the pre-trained style conversion model into the pre-trained discriminant model to obtain the corresponding discriminant result.
[0149] Optionally, the style conversion model includes a word embedding model. When the model training module 120 fixes the model parameters of the pre-trained discriminant model, it can be specifically used for:
[0150] Fix the model parameters of the discriminant model after pre-training and the parameters of the word embedding model.
[0151] It can be understood that each module of the training device provided in the embodiments of the present application may have the function of implementing the corresponding steps in the training method provided in the embodiments of the present application. Among them, this function can be implemented by hardware or by hardware executing corresponding software. The above-mentioned modules can be software and / or hardware, and each module can be implemented separately or multiple modules can be integrated to implement. For the function description of each module of the training device, reference can be specifically made to the corresponding description in the training method in the above-mentioned embodiments, which will not be elaborated here.
[0152] Based on the principle of the training method provided in the embodiments of the present application, the embodiments of the present application also provide a text processing method, which may include:
[0153] Obtain the text to be processed;
[0154] Input the text to be processed into the language style conversion model to obtain a target text with a target language style corresponding to the text to be processed, where the language style conversion model is trained based on the training method of the language style conversion model provided in any embodiment of the present application.
[0155] Optionally, if the language style conversion model is trained based on Figure 3 the training method corresponding to the example shown in, before inputting the text to be processed into the style conversion model, it further includes:
[0156] Label the text to be processed, and the label is used to characterize that the language style of the target text corresponding to the text to be processed is the target language style.
[0157] That is to say, when training the style conversion model by labeling the training sample data with a label for characterizing whether the target text corresponding to the source text is a text with a target language style, when processing the text to be processed based on the trained model, in order to make the output tend to be a text with a target language style, the text to be processed can also be labeled with a label, and this label is the label corresponding to the target language style, such as labeling the label "_p" at the end of the text to be processed.
[0158] Based on the principle of the training method provided in the embodiments of the present application, the embodiments of the present application also provide a text processing device, which may include a module for obtaining the text to be processed and a target text acquisition model. Among them:
[0159] The module for obtaining the text to be processed is used to obtain the text to be processed;
[0160] A target text acquisition model is used to input the text to be processed into a language style conversion model to obtain a target text with a target language style corresponding to the text to be processed, where the language style conversion model is trained based on the training method of the language style conversion model provided in any embodiment of the present application.
[0161] Optionally, if the language style conversion model is trained based on Figure 3 the training method corresponding to the example shown in
[0162] Before inputting the text to be processed into the style conversion model, label the text to be processed, and the label is used to represent that the language style of the target text corresponding to the text to be processed is the target language style.
[0163] Based on the same principle as the solutions described above in this application, an embodiment of the present application also provides a text processing method, which may include:
[0164] Obtain the text to be processed;
[0165] Label the text to be processed, and the label is used to represent that the language style of the target text corresponding to the text to be processed is the target language style;
[0166] Input the text to be processed into the language style conversion model to obtain the target text corresponding to the text to be processed.
[0167] Optionally, the target language style includes but is not limited to honorific style or non - honorific style.
[0168] Correspondingly, an embodiment of the present application also provides a text processing device, which may include:
[0169] A text - to - be - processed acquisition module, used to obtain the text to be processed;
[0170] A text annotation module, used to label the text to be processed, and the label is used to represent that the language style of the target text corresponding to the text to be processed is the target language style;
[0171] A target text acquisition model, used to input the text to be processed into the language style conversion model to obtain the target text corresponding to the text to be processed.
[0172] Based on the same principle as the above - provided methods and devices in this application, an embodiment of the present application also provides an electronic device, which may include a memory and a processor; wherein, a computer program is stored in the memory; the processor is used to call the computer program to execute the method provided in any embodiment of the present application or the steps executed by any device.
[0173] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the methods provided in any embodiment of the present application or the steps executed by any device are implemented.
[0174] As an optional solution, Figure 9 FIG. shows a schematic structural diagram of an electronic device applicable to the embodiments of the present application, as Figure 9 shown. The electronic device 4000 may include a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0175] The processor 4001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in connection with the disclosure of the present application. The processor 4001 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0176] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI bus or an EISA bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is shown in FIG., but it does not mean that there is only one bus or one type of bus.
[0177] The memory 4003 may be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or an EEPROM, a CD-ROM or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage medium or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0178] The memory 4003 is used to store the computer program for executing the solution of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the content shown in any of the foregoing method embodiments.
[0179] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this text, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0180] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A text processing method, characterized in that, including: Obtain the text to be processed; Annotate the label of the text to be processed, where the label is used to characterize that the language style of the target text corresponding to the text to be processed is the target language style; Input the text to be processed with the label into the language style conversion model to obtain the target text in the target language style corresponding to the text to be processed.
2. The method according to claim 1, wherein The target language style is the honorific style or the non - honorific style.
3. A training method for a language style conversion model, characterized in that, including: Obtain training sample data, where the training sample data includes first training texts and second training samples. Each first training text includes a source text in the original language style and a target text in the target language style corresponding to the source text. Each second training sample includes a source text in the original language style and a target text in a non - target language style corresponding to the source text; Train the language style conversion model based on the training sample data until the total loss function of the style conversion model converges. Among them, the total loss function includes a text processing loss function, and the text processing loss function is used to characterize the difference between the text output by the style conversion model and the corresponding target text; The training of the language style conversion model based on the training sample data includes: Annotate the first label of the source text in the training sample data, where the first label is used to characterize whether the language style of the source text is the target language style; Train the style conversion model based on the labeled source text and the corresponding target text.
4. The method according to claim 3, characterized in that, The training of the language style conversion model based on the training sample data further includes: Set up a language style discriminant model, where the discriminant model is used to determine the probability that the text output by the style conversion model is a text with the target language style. Among them, the total loss function further includes a first discriminant loss function corresponding to the discriminant model; Train the style conversion model based on the training sample data and the discriminant model.
5. The method according to claim 4, wherein The training of the language style conversion model based on the training sample data until the total loss function of the style conversion model converges includes: Pre - train the style conversion model based on the training sample data until the text processing loss function converges; Train the pre - trained style conversion model based on the training sample data until the total loss function converges.
6. The method according to claim 5, wherein The training of the language style conversion model based on the training sample data until the total loss function of the style conversion model converges includes: Pre - train the discriminant model based on the output of the pre - trained style conversion model until the second discriminant loss function converges; Train the pre - trained style conversion model based on the training sample data and the pre - trained discriminant model until the total loss function converges.
7. A text processing device, characterized in that, including: A text - to - be - processed acquisition module, used to obtain the text to be processed; A text annotation module, used to annotate the label of the text to be processed, where the label is used to characterize that the language style of the target text corresponding to the text to be processed is the target language style; A target text acquisition model is used to input the to-be-processed text with labels into a language style conversion model to obtain a target text in the target language style corresponding to the to-be-processed text.
8. A training device for a language style conversion model, characterized in that, It includes: A training sample acquisition module, which is used to acquire training sample data. The training sample data includes first training texts and second training samples. Each first training text includes a source text in the original language style and a target text in the target language style corresponding to the source text. Each second training sample includes a source text in the original language style and a target text in a non-target language style corresponding to the source text; A model training module is used to train the language style conversion model based on the training sample data until the total loss function of the style conversion model converges. Among them, the total loss function includes a text processing loss function, and the text processing loss function is used to characterize the difference between the text output by the style conversion model and the corresponding target text; When the model training module trains the language style conversion model based on the training sample data, it is specifically used for: Annotating a first label for the source text in the training sample data, and the first label is used to characterize whether the language style of the source text is the target language style; Training the style conversion model based on the labeled source text and the corresponding target text.
9. An electronic device, characterized in that, It includes a memory and a processor; A computer program is stored in the memory; The processor is used to call the computer program to execute the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Text style conversion method and device, storage medium and computer device
CN109635253A
Machine translation method and device, electronic equipment and medium
CN109697292A