A text style transfer method based on generative adversarial networks

By adopting a generative adversarial network method in text style transfer, using text convolutional network and gated loop unit model for style transfer, and using discriminant model for style judgment feedback, the problems of low style judgment accuracy and text generation quality in the existing technology are solved, and efficient and diverse text style transfer effects are achieved.

CN119578381BActive Publication Date: 2025-06-24NORTHWESTERN POLYTECHNICAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510138731.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-24
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

In the text style transfer, the prior art has problems with low accuracy in sentence style judgment, text generation quality and authenticity, and its dependence on parallel corpus is too high, making it difficult to generate diversified output.

Method used

A generative adversarial network is used to transfer text styles through text convolutional networks, generators and discriminators. The generator includes an encoder and a decoder, which adopts a gated cyclic unit model and a bidirectional self-attention mechanism. This method generates target style statements by deleting attribute words and using discriminant models to perform style judgment feedback.

Benefits of technology

It effectively improves the accuracy of sentence style judgment, the quality and authenticity of text generation, can generate diverse outputs, and shows good robustness on different data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119578381B_ABST
    Figure CN119578381B_ABST
Patent Text Reader

Abstract

The present invention discloses a text style transfer method based on a generative adversarial network, which obtains the predicted attribute style probability of the original sentence through a text convolutional network; deletes the attribute words in the original sentence, and inputs the sentence after deletion into the text convolutional network again until the predicted attribute style probability after deleting each attribute word is obtained; performs a probability difference operation on the predicted attribute style probabilities, arranges each attribute word in descending order according to the result to obtain a sorting result; deletes the first attribute word in the original sentence in descending order to obtain a first sentence; when the attribute style probability and length of the first sentence do not meet the preset conditions, continue to delete the second attribute word on the first sentence until the attribute style probability and length of the sentence after deletion meet the preset conditions to obtain the semantic content; and obtains the target style sentence through an encoder and a decoder for the semantic content. The present invention improves the accuracy of sentence style judgment, the quality and authenticity of text generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text style transfer research, and relates to, but is not limited to, a text style transfer method based on a generative adversarial network. Background Art

[0002] Style is an important part of effective communication. By changing it, we can convey more or less information than the literal meaning of the words. Therefore, the ability to reason about style in language is a fundamental ability required to build successful natural language systems. For example, a children's companion robot replies in a juvenile style, and an elderly care robot replies in an elderly tone. In addition, in professional field products, users hope that the products have a specific style, such as a tone biased towards Shaanxi dialect, containing positive or negative emotions, etc. The goal of text style transfer is to change the style or attributes of the text, such as emotion, gender, etc., by editing or generating while retaining the content text unrelated to the style in the original text.

[0003] In related technologies, text style transfer is achieved by modifying sentences in a continuous space using gradients during the inference process, or by dynamically using the relevance of each output word to the target style for unsupervised style transfer. However, the above methods have problems such as being too fixed in the representation of a specific style, being unable to produce diverse outputs, relying too much on parallel corpora, and having weak discriminative ability for sentences with implicit style expressions.

[0004] Therefore, how to improve the accuracy of sentence style judgment, text generation quality and authenticity has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a text style transfer method based on a generative adversarial network, which at least solves the problems such as low accuracy of sentence style judgment, text generation quality and authenticity in related technologies.

[0006] According to the first aspect of the embodiment of the present invention, there is provided a text style transfer method based on a generative adversarial network, where the generative adversarial network includes a text convolutional network, a generator and a discriminator; the generator includes an encoder and a decoder; the encoder and the decoder include a gated recurrent unit model and a bidirectional self-attention mechanism, including:

[0007] Input the original sentence into the text convolutional network to obtain the predicted attribute style probability;

[0008] Delete the attribute words in the original sentence respectively, and input the sentence after deletion into the text convolutional network again until the predicted attribute style probability corresponding to each deleted attribute word is obtained;

[0009] Perform a probability difference operation between the predicted attribute style probabilities to obtain an operation result; and sort each of the attribute words in descending order based on the operation result and filter out stop words to obtain a sorting result;

[0010] Delete the first attribute word in the original statement in descending order in the sorting result to obtain a first statement;

[0011] When the attribute style probability and length of the first statement do not meet the preset conditions, continue to delete the second attribute word from the first statement until the attribute style probability and length of the statement after deletion meet the preset conditions, then stop deleting and obtain the semantic content of the original statement; the preset conditions are that the attribute style probability of the statement is not greater than a preset probability value or the length of the statement is not greater than a preset length value;

[0012] Map the semantic content through the encoder to obtain a text content representation, and obtain a target style statement based on the text content representation, the target attribute style, and the decoder.

[0013] According to the second aspect of the embodiments of the present invention, there is provided an electronic device, including: a processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete mutual communication through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method described in the first aspect.

[0014] According to the third aspect of the embodiments of the present invention, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.

[0015] According to the solution provided by the embodiments of the present invention, the generative adversarial network includes a text convolutional network, a generator, and a discriminator; the generator includes an encoder and a decoder; the encoder and the decoder include a gated recurrent unit model and a bidirectional self-attention mechanism. The original statement is input into the text convolutional network to obtain the predicted attribute style probability; the attribute words are deleted respectively in the original statement, and the statement after deletion is input into the text convolutional network again until the predicted attribute style probabilities corresponding to each deleted attribute word are obtained; probability difference operations are performed between the predicted attribute style probabilities to obtain an operation result; and each attribute word is sorted in descending order and stop words are filtered according to the operation result to obtain a sorting result; the first attribute word is deleted in the original statement in descending order in the sorting result to obtain a first statement; when the attribute style probability and length of the first statement do not meet the preset conditions, the second attribute word is continuously deleted on the first statement until the attribute style probability and length of the statement after deletion meet the preset conditions, then the deletion stops and the semantic content of the original statement is obtained; the preset conditions are that the attribute style probability of the statement is not greater than the preset probability value or the length of the statement is not greater than the preset length value; the semantic content is mapped through the encoder to obtain a text content representation, and a target style statement is obtained based on the text content representation, the target attribute style, and the decoder. In this process, the style attributes in the original text are removed and deleted, and the statement after deletion is input into the generative adversarial network, and a discriminant model is used to perform style judgment feedback when the model generates a target style statement, so that the style accuracy of the generated target style statement is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0017] Figure 1 is a flowchart showing a text style transfer method based on a generative adversarial network provided by an embodiment of the present invention Figure 1 ;

[0018] Figure 2 is a flowchart showing a text style transfer method based on a generative adversarial network provided by an embodiment of the present invention Figure 2 ;

[0019] Figure 3 is a flowchart showing a text style transfer method based on a generative adversarial network provided by an embodiment of the present invention Figure 3 ;

[0020] Figure 4 A flowchart of a discriminator provided by an embodiment of the present invention;

[0021] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0023] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0024] It should be noted that the terms "first / second / third" involved in the embodiments of the present invention are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.

[0025] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as that generally understood by those of ordinary skill in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0026] Figure 1 A flowchart of a text style transfer method based on a generative adversarial network provided by an embodiment of the present invention Figure 1 A text style transfer method based on a generative adversarial network provided by an embodiment of the present invention can be executed by an electronic device, and the electronic device can be, for example, a computer, a server, etc.

[0027] As Figure 1 shown, the text style transfer method based on a generative adversarial network includes:

[0028] S101. Input the original sentence into the text convolutional network to obtain the predicted attribute style probability.

[0029] In an embodiment of the present invention, a generative adversarial network is proposed for text style transfer. Among them, the generative adversarial network includes a text convolutional network, a generator, and a discriminator; the generator includes an encoder and a decoder; the encoder and decoder include a gated recurrent unit model and a bidirectional self-attention mechanism. First, the text convolutional network is pre-trained according to the attribute text dataset of the non-parallel corpus; the task of training the text convolutional network can be converted into a text classification task. Input the original sentence into the trained text convolutional network to obtain the prediction probability for each predefined style attribute. Among them, the text convolutional network can be a TextCNN network. When using the TextCNN network to classify the attribute style of the input original sentence, the local features of the original sentence are sliding windows composed of several words, similar to N-gram; the text convolutional network can automatically combine and filter the N-gram features. As shown in the following formula (1):

[0030] (1);

[0031] In the above formula (1), is the predicted attribute style probability obtained through the text convolutional network; is the attribute style label of the original sentence, P represents probability, is the parameter set of the text convolutional network, and x is the original sentence.

[0032] S102. Delete the attribute words in the original sentence respectively, and input the sentence after deletion into the text convolutional network again until the predicted attribute style probability corresponding to the deletion of each attribute word is obtained.

[0033] In an embodiment of the present invention, first delete an attribute word in the original sentence, input the sentence after deleting an attribute word into the text convolutional network to obtain the predicted attribute style probability; then re-delete another attribute word in the original sentence, and input the sentence after deleting the attribute word into the text convolutional network again to obtain the predicted attribute style probability. Repeat this process until the predicted attribute style probability corresponding to the deletion of each attribute word is obtained.

[0034] In an embodiment of the present invention, the original sentence in which a certain attribute word (i represents the position of the attribute word w in the sentence) is deleted to obtain , then the attribute style probability after deletion becomes:

[0035] (2);

[0036] In the above formula (2), is the attribute style probability after deleting an attribute word, is the statement after deleting a certain attribute word.

[0037] S103. Perform a probability difference operation on the predicted attribute style probabilities to obtain an operation result; and sort the attribute words in descending order according to the operation result and filter out the stop words to obtain a sorting result.

[0038] In the embodiment of the present invention, the difference before and after deleting the attribute word is calculated through the predicted attribute style probability to obtain an operation result, as shown in the following formula (3):

[0039] (3);

[0040] In the above formula (3), is the probability that the word is an attribute word.

[0041] Exemplarily, when i is 1, find the position of the attribute word in the original statement, obtain the attribute style probability after deleting the attribute word in the original statement, and subtract from the attribute style probability corresponding to the original statement to obtain the operation result ; when i is 2, find the position of the attribute word in the original statement, obtain the attribute style probability after deleting the attribute word in the original statement, and subtract from the attribute style probability corresponding to the original statement to obtain the operation result , and so on.

[0042] Furthermore, sort each attribute word in descending order according to the calculated value, and filter out the stop words to prevent grammar destruction to obtain a sorting result.

[0043] S104. Delete the first attribute word in the original statement in descending order in the sorting result to obtain a first statement.

[0044] In the embodiment of the present invention, the first attribute word is the attribute word corresponding to the maximum probability difference in the sorting result, and the first attribute word is deleted in the original statement to obtain a first statement.

[0045] S105. When both the attribute style probability and the length corresponding to the first statement do not meet the preset conditions, continue to delete the second attribute word from the first statement until the attribute style probability and the length corresponding to the statement after deletion meet the preset conditions, then stop deleting and obtain the semantic content of the original statement; the preset conditions are that the attribute style probability of the statement is not greater than the preset probability value or the length of the statement is not greater than the preset length value.

[0046] In an embodiment of the present invention, when the attribute style probability corresponding to the first statement is not greater than the preset probability value or the length corresponding to the first statement is not greater than the preset length value, stop deleting and obtain the semantic content of the original statement. When the attribute style probability corresponding to the first statement is less than the preset probability value and the length corresponding to the first statement is less than the preset length value, continue to delete the second attribute word (the second attribute word in the sorting result) from the first statement, and repeat the comparison process with the preset conditions until stop deleting and obtain the semantic content of the original statement.

[0047] Further, when the statement after deleting the style word satisfies that the corresponding attribute style probability is less than or the length of the remaining statement is less than When one of the two conditions is met, the deletion process terminates and the semantic content is obtained. The preset conditions are as shown in (4) and (5) below:

[0048] (4);

[0049] (5);

[0050] In the above formula, is a hyperparameter, is the length size of the first statement; is the number of deleted attribute words; is a hyperparameter that controls the content preservation degree of the original statement.

[0051] S106. Map the semantic content through an encoder to obtain a text content representation, and based on the text content representation, the target attribute style, and a decoder, obtain a target style statement.

[0052] In an embodiment of the present invention, map the semantic content to a hidden layer space through an encoder to obtain a text content representation, as shown in the following formula (6):

[0053] (6);

[0054] In the above formula (6), represents the encoder parameter; represents the text feature encoding. The semantic content can be split into m semantic contents, is the first semantic content, For the second semantic content, and so on, C represents the text content representation.

[0055] Further, the text content representation and the target attribute style corresponding to the preset target style statement are input into the decoder to obtain the target style statement.

[0056] Such as Figure 2 shown, Figure 2 is a schematic flow chart of a text style transfer method based on a generative adversarial network provided by an embodiment of the present invention. Figure 2 In Figure 2 it, the application process and training process of the adversarial generative network are included. Figure 2 The overall framework in

[0057] Such as Figure 3 shown, Figure 3 is a schematic flow chart of a text style transfer method based on a generative adversarial network provided by an embodiment of the present invention. Figure 3 In Figure 3 it, style location deletion, an encoder, and a decoder are included. The original statement (from w1 to w n ) is input into the text convolutional network in the generator to obtain the attribute style probability p. Based on the original statement, the attribute words are deleted to obtain the statement after deletion (from c1 to c m ). The statement after deletion is input into the text convolutional network to obtain the attribute style probability . Based on the attribute style probability p and the attribute style probability , the semantic content is obtained. The semantic content is input into the encoder to obtain the text content representation c. The text content representation and the target attribute style s r are input into the decoder to obtain the target style statement B (from y1 to y n ). Among them, both the decoder and the encoder include GRU.

[0058] It can be understood that, in the embodiments of the present invention, the generative adversarial network includes a text convolutional network, a generator, and a discriminator; the generator includes an encoder and a decoder; the encoder and the decoder include a gated recurrent unit model and a bidirectional self-attention mechanism. The original statement is input into the text convolutional network to obtain the predicted attribute style probability; the attribute words are respectively deleted in the original statement, and the statement after deletion is input into the text convolutional network again until the predicted attribute style probabilities corresponding to each deleted attribute word are obtained; the probability difference operation is performed between the predicted attribute style probabilities to obtain an operation result; and each attribute word is sorted in descending order and stop words are filtered according to the operation result to obtain a sorting result; the first attribute word is deleted in the original statement in descending order in the sorting result to obtain a first statement; when the attribute style probability and length of the first statement do not meet the preset conditions, continue to delete the second attribute word on the first statement until the attribute style probability and length of the statement after deletion meet the preset conditions, then the deletion stops and the semantic content of the original statement is obtained; the preset conditions are that the attribute style probability of the statement is not greater than the preset probability value or the length of the statement is not greater than the preset length value; the semantic content is mapped through the encoder to obtain a text content representation, and a target style statement is obtained based on the text content representation, the target attribute style, and the decoder. In this process, the style attributes in the original text are removed and deleted, and the statement after deletion is input into the generative adversarial network, and a discriminant model is used to perform style judgment feedback when the model generates a target style statement, so that the style accuracy of the generated target style statement is effectively improved.

[0059] In some embodiments of the present invention, S101 can be implemented through S1011 to S1012, and the following steps are used for illustration.

[0060] S101. Preprocess the original statement to obtain a preprocessed statement, and use the filter in the text convolutional network to perform local feature extraction on the preprocessed statement to obtain a feature map.

[0061] In some embodiments of the present invention, the original statement is subjected to data preprocessing such as word segmentation and stop word removal to obtain a preprocessed statement, and the preprocessed statement is input into the text convolutional network. The filter in the text convolutional network is used to perform local feature extraction processing on the preprocessed statement, and a corresponding feature map is obtained after the convolution operation.

[0062] S102. Based on the feature map, the pooling layer, the fully connected layer, and the softmax layer in the text convolutional network, obtain the predicted attribute style probability.

[0063] In an embodiment of the present invention, the text convolutional network further includes a pooling layer, a fully connected layer, and a softmax layer. The obtained feature map is input into the pooling layer network for low-dimensional compression of data, and the maximum pooling method is used to obtain local optimal features. Finally, the predicted attribute style probability is output through a fully connected layer and a softmax layer.

[0064] In some embodiments of the present invention, the discriminator includes a bidirectional encoder representation network and a text convolutional network. The training process of the generator and the discriminator can be implemented through the following S201 to S207, and is described through the following steps.

[0065] S201: Input the training original statement sample into the text convolutional network to obtain the predicted first attribute style probability.

[0066] S202: Delete the training attribute words in the training original statement sample respectively, and input the statement sample after deletion into the generator to be trained again until the predicted first attribute style probability corresponding to each deleted training attribute word is obtained.

[0067] S203: Perform a probability difference operation between the predicted first attribute style probabilities to obtain a first operation result; and according to the first operation result, arrange each training attribute word in descending order and filter out stop words to obtain a first sorting result.

[0068] S204: Delete the first training attribute word in the training original statement sample in descending order in the first sorting result to obtain a first statement sample.

[0069] S205: When the first attribute style probability and the first length corresponding to the first statement sample do not meet the preset conditions, continue to delete the second training attribute word on the first statement sample until the first attribute style probability and the first length corresponding to the statement sample after deletion meet the preset conditions, then the deletion stops and the first semantic content of the training original statement sample is obtained; the preset conditions are that the attribute style probability of the statement is not greater than the preset probability value or the length of the statement is not greater than the preset length value.

[0070] S206: Map the first semantic content through an encoder to obtain a first text content representation, and based on the first text content representation, the first target attribute style, and a decoder, obtain a first style statement sample.

[0071] S207: Obtain the trained generator and discriminator based on the first style statement sample.

[0072] In some embodiments of the present invention, the collected training original statement samples are subjected to the above steps S201 to S206 to obtain the first style statement samples. The process is of the same type as that of S201 to S206 and will not be explained here. Based on the first semantic content, the first target attribute style, and the first style statement samples, the trained generator and discriminator are obtained.

[0073] In some embodiments of the present invention, S207 can be implemented through S2071 to S2073 and will be described through the following steps.

[0074] S2071. Obtain the reconstructed training original statement samples through the first semantic content and the first target attribute style.

[0075] S2072. Obtain the reconstruction loss based on the reconstructed training original statement samples, the first semantic content, the first target attribute style, and the reconstruction loss function.

[0076] In some embodiments of the present invention, statement reconstruction is performed on the first semantic content and the corresponding first target attribute style obtained by deleting the attribute style to obtain the reconstructed training original statement samples. The reconstruction loss is obtained based on the reconstructed training original statement samples, the first semantic content, the first target attribute style, and the reconstruction loss function. Among them, the reconstruction loss function is shown as follows (7):

[0077] (7);

[0078] In the above formula (7), c is the semantic content, is the first target attribute style, x is the reconstructed training original statement sample, is the reconstruction loss, are the encoder parameters, and D is the set of reconstructed training original statement samples.

[0079] S2073. Obtain the trained generator and discriminator based on the reconstruction loss and the first style statement samples.

[0080] In some embodiments of the present invention, based on the reconstruction loss and the first style statement samples, the generator and discriminator to be trained are further trained to obtain the trained generator and discriminator.

[0081] In some embodiments of the present invention, S2073 can be implemented through S301 to S305 and will be described through the following steps.

[0082] S301. Use each encoding layer to extract features from the first style statement samples of the input layer, and splice the first CLS vectors of each encoding layer after feature extraction to obtain the spliced feature vectors.

[0083] In an embodiment of the present invention, first, a Bidirectional Encoder Representations from Transformers (BERT network) is used to extract features from the first style statement samples. The structure of the BERT network has 12 encoding layers after the input layer, and the first CLS vector of each encoding layer can be used as the output vector. In the encoder of the BERT network, there is a hierarchical structure between the depth of the layer and the modeling ability of the corresponding statement information. The statement representations in the shallow encoder of the model tend to be low-level semantic information, while as the number of layers deepens, the statement representations in the subsequent layers can better represent high-level semantic information. For each encoding layer of the BERT network, the hidden layer state vector corresponding to the CLS token is extracted therefrom, and the CLS vectors extracted from each encoding layer are concatenated together, and more flexible feature fusion is obtained through a fully connected layer to obtain the concatenated feature vector.

[0084] S302. Input the concatenated feature vector into a text convolutional network to obtain the probability that the first style statement sample is a true statement.

[0085] In some embodiments of the present invention, the concatenated feature vector is used as the input of the text convolutional network. The convolutional layer and pooling layer of the text convolutional network can capture local features in the text and reduce the feature dimension. The output of the convolutional layer passes through the pooling layer and then is connected to a fully connected layer to obtain the probability that the first style statement sample is a true statement. Among them, the true statement is the target expected style statement corresponding to the training original statement sample.

[0086] In some embodiments of the present invention, as Figure 4 shown, Figure 4 is a schematic flowchart of a discriminator provided by an embodiment of the present invention. The BERT network in the discriminator includes 12 encoders. The first style statement sample is input into the 12 encoders. Each encoder outputs a CLS, and each CLS is concatenated to obtain the concatenated feature vector. The concatenated feature vector is input into the text convolutional network in the discriminator network. The text convolutional network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer to obtain a classification label, that is, the probability that the first style statement sample is a true statement is obtained through the fully connected layer.

[0087] S303. Obtain a first attribute classification loss based on the probability of the true statement, and obtain a second attribute classification loss based on the probability of the first target attribute style corresponding to the first style statement sample.

[0088] In some embodiments of the present invention, a first attribute classification loss is calculated based on a first attribute classification loss function and the probability of a true statement, and a second attribute classification loss is calculated based on a second attribute classification loss function and the probability of a first target attribute style corresponding to a first style statement sample. Among them, the first attribute classification loss function and the second attribute classification loss function are respectively shown as follows (8) and (9):

[0089] (8);

[0090] (9);

[0091] In the above formulas (8) and (9), is the attribute style of the true statement ; is the probability that the true statement has the attribute style ; is the first target attribute style; is the probability that the first style statement sample has the first target attribute style ; is the first attribute classification loss, is the second attribute classification loss, and are respectively calculation hyperparameters.

[0092] S304. Adjust the parameters of the generator to be trained based on the second attribute classification loss and the reconstruction loss until the generator is obtained.

[0093] S305. Obtain the trained discriminator based on the second attribute classification loss.

[0094] In some embodiments of the present invention, the parameters of the generator to be trained are adjusted through the second attribute classification loss and the reconstruction loss, and the parameters of the discriminator to be trained are adjusted based on the second attribute classification loss until the training conditions are met, and the trained generator and discriminator are obtained.

[0095] In some embodiments of the present invention, S305 can be implemented through S3051 to S3054, and the following steps are used for illustration.

[0096] S3051. Randomly select a reference statement set and a comparison statement set from the true statements; and calculate the text similarity between the first style statement sample and the reference statement set through the cosine theorem calculation formula.

[0097] In some embodiments of the present invention, during the model learning process, a group of true statements are respectively extracted to construct a reference statement set and a comparison statement set Calculate the text similarity between the first style statement sample and the reference statement set through the cosine theorem calculation formula. The cosine theorem calculation formula (10) is as follows:

[0098] (10);

[0099] In the above formula, is the feature vector of the first style statement sample; is the feature vector of the reference statement set. i is the subscript index of the feature vector, y is the first style statement sample, u is the reference statement, represents the i-th feature vector in the generated statement, represents the i-th feature vector in the reference statement.

[0100] S3052. Calculate the ranking score between the first style statement sample and the comparison statement set through the ranking score calculation formula.

[0101] In some embodiments of the present invention, the ranking score calculation formula is as shown in the following (11). This ranking score calculation formula can be the Softmax formula, and the ranking score between the first style statement sample and the comparison statement set is calculated through the following formula.

[0102] (11);

[0103] In the above formula, is the ranking score calculation parameter, is the comparison statement set, is the sorted first style statement sample, is the sorted comparison statement set.

[0104] S3053. Obtain the expected ranking score corresponding to the first style statement sample based on the ranking score; and obtain the ranking loss based on the expected ranking score.

[0105] In some embodiments of the present invention, the expected ranking score corresponding to the first style statement sample is calculated through the expected ranking score calculation formula and the ranking score. The expected ranking score calculation formula is as shown in the following formula (12):

[0106] (12);

[0107] In the above formula, is the expected ranking score, and U is the reference statement set.

[0108] Furthermore, the ranking loss is calculated through the ranking loss calculation formula and the expected ranking score. The ranking loss calculation formula is as follows (13):

[0109] (13);

[0110] In the above formula (13), is the discriminator parameter, is the ranking loss.

[0111] S3054. Determine the total loss through the ranking loss and the first attribute classification loss; and adjust the parameters of the discriminator through the total loss until the trained discriminator is obtained.

[0112] In some embodiments of the present invention, the total loss calculation formula is as shown in (14) below:

[0113] (14);

[0114] In the above formula (14), is the total loss.

[0115] Furthermore, after determining the total loss through the ranking loss and the first attribute classification loss, adjust the parameters of the discriminator through the total loss until the trained discriminator is obtained. Among them, the loss of the ranking score is taken as negative because the goal of the discriminator is to make the ranking score of the generated sentence as low as possible. In addition, according to the cosine distance between the generated sentence and the reference sentence and the ranking score

[0116] In an embodiment of the present invention, in order to verify the effectiveness of a text style transfer method based on a generative adversarial network provided by the present application, the performance of the present invention in generating text is compared with that of other technical models. The results are shown in Tables 1, 2, and 3. Among them, Table 1 includes a text style transfer method based on an adversarial network and domain adaptation (text styletransfer with adversarial network and, ANDA), back-translation for style transfer (BST), style transfer accuracy (Accuracy, ACC), text retention (Bilingual Evaluation Understudy, BLEU), perplexity (Perplexity, PPL), context-aware style learning and content recovery (Context-aware Style Learning and Content Recove, CSLCR), adversarially regularized autoencoder (Adversarially Regularized Auto, ARAE), delete and retrieve (Delete andRetrieve, DAR), cross-aligned auto-encoder (Cross-aligned Auto-encode, CAAE), style embedding (Style-embedding, SE), blind generative style transform (Blind Generative Style Transform, B-GST), and a text style transfer model based on a generative adversarial network (Text Style Transfer Based on GenerativeAdversarial Networks, TSTG). Among them, TSTG is the model proposed by the present invention, and other models are baseline models.

[0117] Table 1 Performance of Models on the YELP Dataset

[0118] ;

[0119] Table 2 Performance of Models on the AMAZON Dataset

[0120] ;

[0121] Table 3 Performance of Models on the CAPTIONS Dataset

[0122] ;

[0123] In the embodiments of the present invention, it can be seen from the experimental results of synthesizing three tables that the proposed invention has obvious advantages over other technical methods in terms of the accuracy of style transfer, and also has good robustness to datasets with different language distributions. Its comprehensive performance is better than most baseline models.

[0124] Because the present invention first removes the style attributes in the original text and then inputs it into the model, and uses a discriminant model to provide style judgment feedback when the model generates text statements in the target style, the accuracy of the generated text style of the present invention has been effectively improved.

[0125] Compared with other models, the ACC indicators of the present invention on the three datasets of YELP, AMAZON, and CAPTIONS have increased by 83.8%, 39.6%, and 39.3% respectively.

[0126] Generally speaking, the model proposed in this chapter has good performance on the three datasets.

[0127] Refer to Figure 5 , which shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. The specific implementation of the electronic device is not limited in the specific embodiments of the present invention.

[0128] As Figure 5 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.

[0129] Among them:

[0130] The processor 502, the communication interface 504, and the memory 506 communicate with each other through the communication bus 508.

[0131] The communication interface 504 is used to communicate with other electronic devices or servers.

[0132] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above method embodiments.

[0133] Specifically, the program 510 may include program code, and the program code includes computer operation instructions.

[0134] The processor 502 may be a central processing unit (CPU), or a specific application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type, such as one or more CPUs; or may be of different types, such as one or more CPUs and one or more ASICs.

[0135] A memory 506 is used to store a program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0136] Specifically, the program 510 may be used to cause the processor 502 to perform the operations corresponding to the methods described in the above method embodiments.

[0137] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and descriptions in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0138] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention may be split into more components / steps, or two or more components / steps or partial operations of components / steps may be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0139] The methods according to the embodiments of the present invention may be implemented in hardware, firmware, or may be implemented as software or computer code stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or may be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and stored in a local recording medium, so that the methods described herein may be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0140] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0141] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention shall be defined by the claims.

Claims

1. A text style transfer method based on generative adversarial network, characterized in that: The generative adversarial network includes a text convolutional network, a generator and a discriminator; the generator includes an encoder and a decoder; the encoder and the decoder include a gated recurrent unit model and a bidirectional self-attention mechanism, the discriminator includes a bidirectional encoder representation network and the text convolutional network, the bidirectional encoder representation network includes an input layer and 12 encoding layers, including: Inputting the original sentence into the text convolutional network to obtain the predicted attribute style probability; Deleting attribute words in the original sentence respectively, and inputting the deleted sentences into the text convolutional network again, until the predicted attribute style probability corresponding to each attribute word after deletion is obtained; Performing a probability difference operation on the predicted attribute style probabilities to obtain an operation result; and arranging each attribute word in descending order and filtering stop words according to the operation result to obtain a sorting result; In the sorting result, the first attribute word is deleted from the original sentence in descending order to obtain a first sentence; When both the attribute style probability and the length corresponding to the first sentence do not meet the preset conditions, continue to delete the second attribute word on the first sentence until the attribute style probability and the length of the deleted sentence meet the preset conditions, then the deletion stops and the semantic content of the original sentence is obtained; the preset conditions are that the attribute style probability of the sentence is not greater than the preset probability value or the length of the sentence is not greater than the preset length value; Mapping the semantic content through the encoder to obtain a text content representation, and obtaining a target style sentence based on the text content representation, the target attribute style and the decoder; The mapping of the semantic content through the encoder to obtain a text content representation, and obtaining a target style sentence based on the text content representation, the target attribute style and the decoder, includes: Mapping the semantic content to a hidden space through the encoder to obtain the text content representation; The target attribute style corresponding to the text content representation and a preset target style sentence is input into the decoder to obtain the target style sentence.

2. The method according to claim 1, characterized in that The inputting the original sentence into the text convolutional network to obtain the predicted attribute style probability includes: Preprocessing the original sentence to obtain a preprocessed sentence, and using the filter in the text convolution network to extract local features of the preprocessed sentence to obtain a feature map; The predicted attribute style probability is obtained based on the feature map, the pooling layer, the fully connected layer and the softmax layer in the text convolutional network.

3. The method according to claim 1, characterized in that The generator and the discriminator are obtained through the following training process: Inputting the original training sentence sample into the text convolutional network to obtain the predicted first attribute style probability; Deleting the training attribute words in the original training sentence samples respectively, and inputting the deleted sentence samples into the generator to be trained again, until the predicted first attribute style probability corresponding to each training attribute word is obtained after deletion; Performing a probability difference operation on the predicted first attribute style probabilities to obtain a first operation result; And according to the first operation result, each training attribute word is arranged in descending order and stop words are filtered out to obtain a first sorting result; In the first sorting result, the first training attribute word is deleted from the original training sentence sample in descending order to obtain a first sentence sample; When the first attribute style probability and the first length corresponding to the first sentence sample do not meet the preset condition, continue to delete the second training attribute word on the first sentence sample until the first attribute style probability and the first length corresponding to the deleted sentence sample meet the preset condition, then the deletion stops and the first semantic content of the original training sentence sample is obtained; The preset condition is that the attribute style probability of the sentence is not greater than a preset probability value or the length of the sentence is not greater than a preset length value; Mapping the first semantic content through the encoder to obtain a first text content representation, and obtaining a first style sentence sample based on the first text content representation, a first target attribute style and the decoder; The trained generator and the discriminator are obtained based on the first semantic content, the first target attribute style and the first style sentence sample.

4. The method according to claim 3, characterized in that The step of acquiring the trained generator and the discriminator based on the first semantic content, the first target attribute style and the first style sentence sample includes: Obtaining a reconstructed training original sentence sample through the first semantic content and the first target attribute style; Obtaining a reconstruction loss based on the reconstructed original training sentence sample, the first semantic content, the first target attribute style, and a reconstruction loss function; The trained generator and the discriminator are obtained based on the reconstruction loss and the first style sentence sample.

5. The method according to claim 4, characterized in that The step of obtaining the trained generator and the discriminator based on the reconstruction loss and the first style sentence sample includes: Using each encoding layer to extract features from the first style sentence sample of the input layer, and concatenating the first CLS vector of each encoding layer after feature extraction to obtain a concatenated feature vector; Inputting the concatenated feature vector into the text convolutional network to obtain the probability that the first style sentence sample is a true sentence; Obtaining a first attribute classification loss based on the probability of the true sentence, and obtaining a second attribute classification loss based on the probability of the first target attribute style corresponding to the first style sentence sample; Adjusting the parameters of the generator to be trained based on the second attribute classification loss and the reconstruction loss until the generator is obtained; The trained discriminator is obtained based on the second attribute classification loss.

6. The method according to claim 5, characterized in that The step of obtaining the trained discriminator based on the second attribute classification loss includes: Randomly selecting from the real sentences to obtain a reference sentence set and a comparison sentence set; and calculating the text similarity between the first style sentence sample and the reference sentence set by using a cosine theorem calculation formula; Calculating the ranking scores of the first style sentence sample and the comparison sentence set by using a ranking score calculation formula; Obtaining an expected ranking score corresponding to the first style sentence sample based on the ranking score; and obtaining a ranking loss based on the expected ranking score; The total loss is determined by the ranking loss and the first attribute classification loss; and the parameters of the discriminator are adjusted by the total loss until a trained discriminator is obtained.

Citation Information

Patent Citations

  • Method for generating ancient poem based on generative adversarial network

    CN114048755A

  • Text generation method based on conditional generative adversarial network syntactic structure control

    CN118153582A

  • Text generation model training method and device

    CN119830001A

  • Keyword generation method and apparatus, and electronic device and computer storage medium

    WO2022134759A1

  • Text analysis method and apparatus, and emotion classification model, medium, terminal and product

    WO2024114382A1