Unsupervised text style migration method and system based on style decoupling
By building a deep learning network model and utilizing style decoupling technology and contrastive learning, we achieve efficient text style transfer without losing semantic content, solving the problem of inaccurate style transfer in existing technologies and improving the stability and accuracy of style conversion.
Patent Information
- Application Number
- CN202510896275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
AI Technical Summary
Existing text style transfer technology is difficult to achieve accurate style conversion without losing semantic content, especially under non-parallel dataset conditions, where the accuracy and stability of style transfer are insufficient.
An unsupervised text style transfer method based on style decoupling is adopted. By constructing a deep learning network model, using a style information extractor and an intermediate text generator, and combining contrastive learning and reconstruction training strategies, the text style is decoupled and a neutral intermediate text is generated to achieve style transfer.
The accuracy and stability of text style transfer are improved, ensuring the integrity of content and the accuracy of style conversion, and adapting to diverse language style requirements.
Smart Images

Figure CN120724980A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and in particular relates to an unsupervised text style transfer method and system based on style decoupling. Background Art
[0002] Text style transfer (TST), a key task in natural language processing (NLP), aims to modify the linguistic style or expression of a text while preserving its original meaning. Advances in natural language processing (NLP) have enabled machines to understand and generate more natural language. However, language inherently possesses a rich set of stylistic attributes, such as formality, politeness, emotional overtones, and personalization. These stylistic factors often determine the adaptability of language in different social contexts. Language usage varies significantly across different contexts, such as between formal business emails and informal conversations with friends, or between news reports and social media comments. This diversity in linguistic style requires natural language processing systems to possess not only semantic understanding capabilities but also the ability to flexibly adapt style. With the widespread adoption of applications such as intelligent dialogue systems, automatic text generation, and intelligent writing assistants, effectively controlling and transferring linguistic style has become crucial for improving the quality and interactive effectiveness of machine-generated language. Due to the complexity and ambiguity of linguistic style, accurately transferring style without losing semantic content remains a major challenge in current natural language processing research. Research on text style transfer not only holds academic value but also demonstrates broad significance and potential in multiple practical applications.
[0003] First, TST technology can help clean up offensive speech online. By transforming text containing biased, insulting, or extreme language into more objective and neutral expressions, it can effectively reduce the spread of online violence, hate speech, and inappropriate language, thereby promoting healthy online communication and social interaction. Second, TST plays a crucial role in enhancing the data diversity of machine learning models. By automatically adjusting text style and generating corpus data in a variety of styles, it can expand training datasets, thereby improving the generalization ability of machine learning models and enabling more accurate predictions and generation across diverse language styles and contexts. Furthermore, TST technology can optimize the human-machine conversation experience by adjusting the robot's tone, style, and emotional expression to better align with users' emotional needs and communication habits. For example, in emotionally supportive conversations, an empathetic language style can enhance the robot's approachability and improve the user's interactive experience. These applications not only improve the quality of machine-generated text but also promote the in-depth development of natural language processing technology in multiple fields, playing a particularly important role in key areas such as social media, intelligent customer service, and content moderation.
[0004] Text style transfer (TST) methods can be divided into those based on parallel datasets and those based on non-parallel datasets. Parallel dataset-based methods rely on well-labeled text pairs with multiple styles and typically employ supervised learning using neural sequence-to-sequence models (such as the Transformer). However, parallel datasets are expensive to construct and struggle to cover diverse styles. Consequently, researchers have focused more on methods based on non-parallel datasets, such as disentanglement and entangled representation methods. Disentangled methods achieve style transfer by separating content and style information, often employing variational autoencoders (VAEs) and generative adversarial networks (GANs). Entangled representation methods, such as adversarial regularized autoencoders (ARAEs) and methods based on external style vectors, directly modify the overall representation of the text. These technological advances have improved the feasibility and flexibility of style transfer.
[0005] In recent years, TST combined with large language models (LLMs) has become a research hotspot, including fine-tuning model parameters and hint-based style transfer. The fine-tuning method introduces style labels into the pre-trained model to reduce training costs and improve transfer quality, while the hint method guides LLMs to generate target style text by designing appropriate inputs. In addition, researchers have also explored prototype-based text editing and pseudo-parallel dataset methods, such as retrieving similar text or using generative models (such as GANs) to generate pseudo data to enhance the effect of style transfer. With the development of deep learning and large models, TST technology is constantly being optimized to achieve more accurate and efficient text style transfer. Summary of the Invention
[0006] The purpose of the present invention is to provide an unsupervised text style transfer method and system based on style decoupling, which are conducive to improving the accuracy and stability of text style transfer.
[0007] To achieve the above objectives, the present invention adopts a technical solution: an unsupervised text style transfer method based on style decoupling, comprising the following steps:
[0008] Step A: Obtain text data of different styles and construct dataset D;
[0009] Step B: Constructing a deep learning network model G; the deep learning network model G passes the input text into a style information extractor for encoding to decouple and extract the style information in the text, and uses contrastive learning to enhance the ability to extract style information; at the same time, the input text is passed into an intermediate text generator to obtain an intermediate text that is different from the input text in terms of text style; then, the outputs of the style information extractor and the intermediate text generator are passed into a style transfer module to perform text style transfer; the deep learning network model G is trained using the dataset;
[0010] Step C: Input the text to be transferred into the trained deep learning network model G and output the text after style transfer.
[0011] Furthermore, the step A specifically includes the following steps:
[0012] Step A1: Obtain a public dataset X of different styles a and public dataset X b , X a With X b The texts in are randomly paired into samples, each of which contains texts with style s a The text t a and style b The text t b ; The data set D consists of multiple samples;
[0013] Step A2: Divide the dataset D into training set D by random sampling n and validation set D v The training set is used for model training and parameter optimization, and the validation set is used to evaluate the performance of the model and perform hyperparameter adjustment.
[0014] Furthermore, the step B specifically includes the following steps:
[0015] Step B1: Set the training set D n The text of each sample t b The input is encoded into a style information extractor to decouple the style information and content information in the text; the style information extractor includes: a style encoder and a feedforward neural network; the input text is first encoded by the style encoder to obtain a feature vector, and then the feature vector is further extracted and mapped by the feedforward neural network to obtain the style representation z b ;
[0016] Step B2: Through comparative learning, the style representations of texts with the same style are made similar, while the style representations of texts with different styles are separated. A style classifier is used to further score the style representations. By improving the score of the style representations, the extracted style representations can accurately represent the text style and provide style information in the subsequent style transfer task. Through training, the style information extractor can effectively capture the style characteristics of the text, providing high-quality style representation vectors for the style transfer task.
[0017] Step B3: Set the training set D n The text of each sample t a The intermediate text generator is input, which uses a text generation model and combines it with a continuous decoding algorithm to generate an intermediate text t that is more neutral in style. n, providing a more neutral text basis for subsequent style transfer;
[0018] Step B4: Train the intermediate text generator through style removal task and semantic preservation task to remove the original style s a While retaining the input text t a The semantic content of
[0019] Step B5: Transform the input text t through the trained intermediate text generator a Convert to style-neutral intermediate text n , then the middle text t n and style representation b Combined, as the input of the style transfer module, to obtain the text after style transfer The style transfer module introduces a text generation model. The style transfer module adopts a reconstruction training strategy to enable the model to learn different style transfer modes while ensuring that the generated text conforms to the target style while maintaining the semantic integrity of the original text.
[0020] Furthermore, the step B1 specifically includes the following steps:
[0021] Step B11: Enter text t b After entering the style information extractor, it first enters the style encoder, which uses the pre-trained model BERT as the encoder model to perform preliminary style feature extraction, that is, the BERT model first performs a pre-trained analysis on the input text t b Processing is performed to obtain the hidden state representation:
[0022] H=BERT(t b )
[0023] in, For input text t b The hidden state matrix, L is the input text t b The length of , d is the hidden layer dimension of the BERT model, which is also the dimension of style representation;
[0024] In order to obtain text-level style features, the output vector corresponding to the [CLS] tag in H is used as the style feature vector:
[0025] h s =H CLS
[0026] Among them, h s ∈R d For input text t b The style feature vector of
[0027] Step B12: Use a feedforward neural network to generate the style feature vector h sPerform feature extraction and mapping to obtain a more stable style representation z:
[0028] s=σ(W1h s +b1)
[0029] z=W2s+b2
[0030] Among them, W1 and W2 are trainable weight matrices in the feedforward neural network, b1 and b2 are bias terms, and σ is a nonlinear activation function. For the final style representation.
[0031] Furthermore, the step B2 specifically includes the following steps:
[0032] Step B21: training the style information extractor through contrastive learning so that the style features of texts with the same style are similar, while the style features of texts with different styles are far apart;
[0033] Define positive samples as texts with the same style, and negative samples as texts with different styles. Maximize the similarity of style representations between positive samples, and minimize the similarity of style representations between negative samples. Assume that the number of texts in each training batch is N, the set of positive samples is J, and the set of negative samples is M. Define the loss function as follows:
[0034]
[0035] Where sim(·,·) represents the cosine similarity function and τ is the temperature coefficient, which is used to control the smoothness or sharpness of the distribution;
[0036] Step B22: Introduce the style classification task and use cross-entropy loss to ensure that the extracted style representation can effectively distinguish different styles:
[0037]
[0038] Among them, C is a style classifier pre-trained on dataset D, s i is the real style category;
[0039] Step B23: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using backpropagation to minimize the loss function to train the style information extractor.
[0040] Furthermore, the specific implementation method of step B3 is:
[0041] Enter the text t a The intermediate text generator is passed in. The intermediate text generator adopts a text generation model and inputs the text t aGo through two processes: encoding and decoding;
[0042] Enter text a First, it is converted into a word vector representation. Each meaningful word is converted into a token and then input into the encoder of the text generation model for encoding:
[0043] H=Encoder(t a )
[0044] in, For input text t a The hidden state matrix, L is the input text t a The length of , d is the hidden layer dimension of the generative model, that is, the feature dimension;
[0045] In the decoding phase, a continuous decoding algorithm is used to generate continuous embedding vectors. At each time step, the decoder calculates the probability distribution of the current token and performs a weighted summation with the word embedding matrix:
[0046] h t =Decoder(v t-1 ,H)
[0047]
[0048] v t =p t W
[0049] Where W is the word embedding matrix with a size of d×K, where K represents the size of the vocabulary; h t is the output of the decoder of the text generation model at time step t, where the value range of t is [1, L], and h i Indicates h t The i-th element of , when t = 1, v0 is the starting symbol vector; p t is the word probability distribution calculated by the softmax function;
[0050] The generated embedding vector v t As input for the next time step, we get the new output:
[0051] h t+1 =Decoder(v t ,H)
[0052] After L time steps, we get P = {p1, p2, ..., p L}, and then generate the intermediate text t n :
[0053] t n =argmax(P)·W
[0054] Among them, the argmax function is used to obtain the word probability distribution p t The index of the word with the highest probability.
[0055] Furthermore, the step B4 specifically includes the following steps:
[0056] Step B41: Train the intermediate text generator through the style removal task;
[0057] s n =C(t n )
[0058]
[0059] in, is the intermediate loss, s n Is the generated text t n The style score of , which takes values of -1 (indicating negative), 1 (indicating positive) or 0 (indicating neutral), and the training goal is to make the style score as close to 0 as possible;
[0060] Step B42: Calculate the original text t a With the middle text t n The cosine similarity between:
[0061]
[0062] in, is the semantic loss, h a and h n The original text t a and the middle text t n The semantic vector obtained by word embedding is trained to minimize the loss and ensure the semantic consistency between the two.
[0063] Step B43: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using backpropagation to minimize the loss function to train the intermediate text generator.
[0064] Furthermore, the step B5 specifically includes the following steps:
[0065] Step B51: Use the style transfer module to transfer the intermediate text t n Style representation z with target style b To process; first, the intermediate text t n Processed by the encoder of the style transfer module to obtain the semantic representation of the text:
[0066] Hn =Encoder(t n ,z b )
[0067] in, For the middle text t n The hidden state matrix, L is the input text t n The length of , d is the hidden layer dimension of the generative model, that is, the feature dimension;
[0068] x t =Decoder(x t-1 ,H n )
[0069]
[0070] Among them, x t-1 Represents the hidden state obtained in the previous time step. When t=1, the hidden state x0 is the starting symbol vector; y t is the predicted probability distribution of the tth word calculated by the Softmax function at the current time step t;
[0071] After L time steps, we get Y = {y1,y2,…,y L}, Jiner generates text after style transfer
[0072]
[0073] Step B52: adopting a reconstruction training strategy during the training of the style transfer module;
[0074] The original text t a Convert to target style text Its losses for:
[0075]
[0076] Among them, G is the style transfer module, ||·|| represents the text similarity measurement, and the goal is to make With the original text a As close as possible to ensure semantic integrity;
[0077] At the same time, to ensure that the generated text conforms to the target style, its loss for:
[0078]
[0079] Among them, s b The target style of the generated sample;
[0080] Step B53: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using back propagation to minimize the loss function to train the model.
[0081] The present invention also provides an unsupervised text style transfer system based on style decoupling, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above-mentioned method steps can be implemented.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] 1) This paper proposes a style disentanglement method that effectively removes style interference by generating style-independent intermediate representations, thereby ensuring content integrity.
[0084] 2) This paper introduces a contrastive learning mechanism to optimize the style information extractor so that it can learn more independent style features, reduce interference with content information, and improve the accuracy of style transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 This is a flow chart for implementing the unsupervised text style transfer method based on style decoupling provided by an embodiment of the present invention;
[0086] Figure 2 It is an architectural diagram of the deep learning network model G in an embodiment of the present invention. DETAILED DESCRIPTION
[0087] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0088] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.
[0089] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0090] like Figure 1 As shown, this embodiment provides an unsupervised text style transfer method based on style decoupling, including the following steps:
[0091] Step A: Obtain text data of different styles and construct dataset D;
[0092] Step B: Constructing a deep learning network model G; the deep learning network model G passes the input text into a style information extractor for encoding to decouple and extract the style information in the text, and uses contrastive learning to enhance the ability to extract style information, making the obtained style information more pure; at the same time, the input text is passed into an intermediate text generator to obtain an intermediate text that is different from the input text in text style; then, the outputs of the style information extractor and the intermediate text generator are passed into a style transfer module to perform text style transfer; the deep learning network model G is trained using the dataset;
[0093] Step C: Input the text to be transferred into the trained deep learning network model G and output the text after style transfer.
[0094] In this embodiment, step A specifically includes the following steps:
[0095] Step A1: We can download public datasets X from the Internet in multiple fields as needed, such as catering, business, and movie reviews. Each public dataset contains a large number of natural language sentences and shares certain specific text features. For example, restaurant reviews may show positive or negative sentiment tendencies. We define these shared features as style, which we represent with the symbol s. Public datasets of different styles contain text with different styles.
[0096] The goal of style transfer is to transfer the style of a sentence from the original style to the target style while maintaining the semantic information of the original text. Here we assume that the original style is s a , the target style is style s b .
[0097] First, download the public dataset X of the corresponding style a and public dataset X b , X a With X b The texts in are randomly paired into samples, each of which contains texts with style s a The text t a and style b The text t b ; The data set D consists of multiple samples.
[0098] Step A2: Divide 80% of the data into training set D by random sampling. n , 20% of the data is divided into the validation set D vThe training set is used for model training and parameter optimization, and the validation set is used to evaluate the performance of the model and perform hyperparameter adjustment.
[0099] Figure 2 This is the architecture diagram of the deep learning network model G in this embodiment. Figure 2 As shown, the specific implementation steps of step B are as follows.
[0100] Step B1: Set the training set D n The text of each sample t b The input is then encoded into a style information extractor to decouple the style information and content information in the text, enabling the model to effectively identify and process different style features. The style information extractor consists of two core components: a style encoder and a feedforward neural network. The input text is first encoded by the style encoder to obtain a feature vector, and then the feature vector is further extracted and mapped by the feedforward neural network to obtain the style representation z. b .
[0101] In this embodiment, step B1 specifically includes the following steps:
[0102] Step B11: Enter text t b After entering the style information extractor, it first enters the style encoder, which uses the pre-trained model BERT as the encoder model to perform preliminary style feature extraction, that is, the BERT model first performs a pre-trained analysis on the input text t b Processing is performed to obtain the hidden state representation:
[0103] H=BERT(t b )
[0104] in, For input text t b The hidden state matrix, L is the input text t b , d is the hidden layer dimension of the BERT model, and is also the dimension of style representation.
[0105] In order to obtain text-level style features, the output vector corresponding to the [CLS] tag in H is used as the style feature vector:
[0106] h s =H CLS
[0107] in, For input text t b The style feature vector of .
[0108] Step B12: Use a feedforward neural network to generate the style feature vector h s Perform feature extraction and mapping to obtain a more stable style representation z:
[0109] s=σ(W1h s +b1)
[0110] z=W2s+b2
[0111] Among them, W1 and W2 are trainable weight matrices in the feedforward neural network, b1 and b2 are bias terms, and σ is a nonlinear activation function. For the final style representation.
[0112] Step B2: Through comparative learning, the style representations of texts with the same style are made similar, while being separated from the style representations of texts with different styles, to enhance the ability to distinguish styles and avoid interference from content information. A style classifier is used to further score the style representations. By improving the score of the style representations, it is ensured that the extracted style representations can accurately represent the text style and provide stable style information in subsequent style transfer tasks. Through training, the style information extractor can effectively capture the style characteristics of the text and provide high-quality style representation vectors for the style transfer task.
[0113] In this embodiment, step B2 specifically includes the following steps:
[0114] Step B21: In order to enhance the distinguishing ability of style features and obtain a more stable style representation, the style information extractor is trained through contrastive learning so that the style features of texts with the same style are similar, while the style features of texts with different styles are far away from each other.
[0115] Define positive samples as texts with the same style, and negative samples as texts with different styles. Maximize the similarity of style representations between positive samples, and minimize the similarity of style representations between negative samples. Assume that the number of texts in each training batch is N, the set of positive samples is J, and the set of negative samples is M. Define the loss function as follows:
[0116]
[0117] Here, sim(·,·) represents the cosine similarity function and τ is the temperature coefficient, which is used to control the smoothness or sharpness of the distribution.
[0118] Step B22: Introduce the style classification task and use cross-entropy loss to ensure that the extracted style representation can effectively distinguish different styles:
[0119]
[0120] Among them, C is a style classifier pre-trained on dataset D, s i Real style category.
[0121] Step B23: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using backpropagation to minimize the loss function to train the style information extractor.
[0122] Step B3: Set the training set D n The text of each sample t a The intermediate text generator is input, which uses a text generation model and combines it with a continuous decoding algorithm to generate an intermediate text t that is more neutral in style. n , providing a more neutral text basis for subsequent style transfer.
[0123] In this embodiment, the specific implementation method of step B3 is:
[0124] Enter the text t a The intermediate text generator is passed in. The intermediate text generator adopts a text generation model and inputs the text t a Go through two processes: encoding and decoding.
[0125] Enter text a First, it is converted into a word vector representation. Each meaningful word is converted into a token and then input into the encoder of the text generation model for encoding:
[0126] H=Encoder(t a )
[0127] in, For input text t a The hidden state matrix, L is the input text t a The length of , d is the hidden layer dimension of the generative model, that is, the feature dimension.
[0128] In the decoding phase, a continuous decoding algorithm is used to generate continuous embedding vectors. At each time step, the decoder calculates the probability distribution of the current token and performs a weighted summation with the word embedding matrix:
[0129] h t =Decoder(v t-1 ,H)
[0130]
[0131] v t =p t W
[0132] Where W is the word embedding matrix with a size of d×K, where K represents the size of the vocabulary; h tis the output of the decoder of the text generation model at time step t, where the value range of t is [1, L], and h i Indicates h t The i-th element of , when t = 1, v0 is the starting symbol vector; p t It is the word probability distribution calculated by the softmax function.
[0133] The generated embedding vector v t As input for the next time step, we get the new output:
[0134] h t+1 =Decoder(v t ,H)
[0135] After L time steps, we get P = {p1, p2, ..., p L}, and then generate the intermediate text t n :
[0136] t n =argmax(P)·W
[0137] Among them, the argmax function is used to obtain the word probability distribution p t The index of the word with the highest probability.
[0138] Step B4: Train the intermediate text generator through style removal tasks and semantic preservation tasks, with the goal of removing the original style as much as possible. a While retaining the input text t a semantic content.
[0139] In this embodiment, step B4 specifically includes the following steps:
[0140] Step B41: To ensure that the intermediate text has a more neutral style, the intermediate text generator is trained through the style removal task;
[0141] s n =C(t n )
[0142]
[0143] in, is the intermediate loss, s n Is the generated text t n The style score is -1 (negative), 1 (positive) or 0 (neutral), and the training goal is to make the style score as close to 0 as possible.
[0144] Through this loss constraint, the model can effectively remove style information during the generation process, ensuring that the generated text is as close to style neutral as possible.
[0145] Step B42: To ensure that the generated intermediate text retains the original semantic information, calculate the original text t a With the middle text t n The cosine similarity between:
[0146]
[0147] in, is the semantic loss, h a and h n The original text t a and the middle text t n The semantic vector obtained by word embedding has the training goal of minimizing the loss and ensuring the semantic consistency between the two.
[0148] Step B43: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using backpropagation to minimize the loss function to train the intermediate text generator.
[0149] Step B5: Transform the input text t through the trained intermediate text generator a Convert to style-neutral intermediate text n , then the middle text t n and style representation b Combined, as the input of the style transfer module, to obtain the text after style transfer The style transfer module introduces a text generation model. To enhance the flexibility and stability of style transfer, the style transfer module adopts a reconstruction training strategy. This enables the model to learn different style transfer modes while ensuring that the generated text conforms to the target style while maintaining the semantic integrity of the original text.
[0150] In this embodiment, step B5 specifically includes the following steps:
[0151] Step B51: After training the intermediate text generator in step B4, the input text t can be stably converted to a Convert to intermediate text n . Through the style transfer module, the intermediate text t n Style representation z with target style b To process; first, the intermediate text t n Processed by the encoder of the style transfer module to obtain the semantic representation of the text:
[0152] H n =Encoder(t n ,z b)
[0153] in, For the middle text t n The hidden state matrix, L is the input text t n The length of , d is the hidden layer dimension of the generative model, that is, the feature dimension;
[0154] x t =Decoder(x t-1 ,H n )
[0155]
[0156] Among them, x t-1 Represents the hidden state obtained in the previous time step. When t=1, the hidden state x0 is the starting symbol vector; y t It is the predicted probability distribution of the tth word calculated by the Softmax function at the current time step t.
[0157] After L time steps, we get Y = {y1,y2,…,y L}, Jiner generates text after style transfer
[0158]
[0159] Step B52: Due to the shortage of data resources and in order to enhance the stability and generalization ability of the style transfer module, a reconstruction training strategy is adopted in the training process of the style transfer module.
[0160] The original text t a Convert to target style text Its losses for:
[0161]
[0162] Among them, G is the style transfer module, ||·|| represents the text similarity measurement, and the goal is to make With the original text a As close as possible to ensure semantic integrity.
[0163] At the same time, to ensure that the generated text conforms to the target style, its loss for:
[0164]
[0165] Among them, s b is the target style of the generated samples.
[0166] Step B53: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using back propagation to minimize the loss function to train the model.
[0167] In this embodiment, the performance of the deep learning network model G provided by this method is compared with other baseline models on the Yelp and IMDB datasets, and the results are shown in Table 1. As can be seen from Table 1, the performance of the model provided by this method is better than that of other models.
[0168] Table 1 Performance comparison of the proposed model and the baseline model on the Yelp and IMDB datasets
[0169]
[0170] This embodiment also provides an unsupervised text style transfer system based on style decoupling, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above-mentioned method steps can be implemented.
[0171] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0172] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0173] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0175] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
Claims
1. An unsupervised text style transfer method based on style decoupling, characterized by: The following steps are involved: Step A: Obtain text data of different styles and construct dataset D; Step B: Constructing a deep learning network model G; the deep learning network model G passes the input text into a style information extractor for encoding to decouple and extract the style information in the text, and uses contrastive learning to enhance the ability to extract style information; at the same time, the input text is passed into an intermediate text generator to obtain an intermediate text that is different from the input text in terms of text style; then, the outputs of the style information extractor and the intermediate text generator are passed into a style transfer module to perform text style transfer; the deep learning network model G is trained using the dataset; Step C: Input the text to be transferred into the trained deep learning network model G and output the text after style transfer.
2. The unsupervised text style transfer method based on style decoupling according to claim 1 is characterized in that The step A specifically comprises the following steps: Step A1: Obtain a public dataset X of different styles a and public dataset X b , X a With X b The texts in are randomly paired into samples, each of which contains texts with style s a The text t a and style v The text t b ; The data set D consists of multiple samples; Step A2: Divide the dataset D into training set D by random sampling n and validation set D v The training set is used for model training and parameter optimization, and the validation set is used to evaluate the performance of the model and perform hyperparameter adjustment.
3. The unsupervised text style transfer method based on style decoupling according to claim 1 is characterized in that The step B specifically comprises the following steps: Step B1: Set the training set D n The text of each sample t b The input is encoded into the style information extractor to decouple the style information and content information in the text; the style information extractor includes: a style encoder and a feedforward neural network; the input text is first encoded by the style encoder to obtain a feature vector, and then the feature vector is further extracted and mapped by the feedforward neural network to obtain the style representation z b ; Step B2: Through comparative learning, the style representations of texts with the same style are made similar, while the style representations of texts with different styles are separated. A style classifier is used to further score the style representations. By improving the score of the style representations, the extracted style representations can accurately represent the text style and provide style information in the subsequent style transfer task. Through training, the style information extractor can effectively capture the style characteristics of the text, providing high-quality style representation vectors for the style transfer task. Step B3: Set the training set D n The text of each sample t a The intermediate text generator is input, which uses a text generation model and combines it with a continuous decoding algorithm to generate an intermediate text t that is more neutral in style. n , providing a more neutral text basis for subsequent style transfer; Step B4: Train the intermediate text generator through style removal task and semantic preservation task to remove the original style s a While retaining the input text t a The semantic content of Step B5: Transform the input text t through the trained intermediate text generator a Convert to style-neutral intermediate text n , then the middle text t n and style representation b Combined, as the input of the style transfer module, to obtain the text after style transfer The style transfer module introduces a text generation model. The style transfer module adopts a reconstruction training strategy to enable the model to learn different style transfer modes while ensuring that the generated text conforms to the target style while maintaining the semantic integrity of the original text.
4. The unsupervised text style transfer method based on style decoupling according to claim 3 is characterized in that The step B1 specifically includes the following steps: Step B11: Enter text t b After entering the style information extractor, it first enters the style encoder, which uses the pre-trained model BERT as the encoder model to perform preliminary style feature extraction, that is, the BERT model first performs a pre-trained analysis on the input text t b Processing is performed to obtain the hidden state representation: H=BERT(t b ) in, For input text t b The hidden state matrix, L is the input text t b The length of , d is the hidden layer dimension of the BERT model, which is also the dimension of style representation; In order to obtain text-level style features, the output vector corresponding to the [CLS] tag in H is used as the style feature vector: h s =H CLS in, For input text t b The style feature vector of Step B12: Use a feedforward neural network to generate the style feature vector h s Perform feature extraction and mapping to obtain a more stable style representation z: s=σ(W1h s +b1) z=W2s+b2 Among them, W1 and W2 are trainable weight matrices in the feedforward neural network, b1 and b2 are bias terms, and σ is a nonlinear activation function. For the final style representation.
5. The unsupervised text style transfer method based on style decoupling according to claim 4 is characterized in that The step B2 specifically includes the following steps: Step B21: training the style information extractor through contrastive learning so that the style features of texts with the same style are similar, while the style features of texts with different styles are far apart; Define positive samples as texts with the same style, and negative samples as texts with different styles. Maximize the similarity of style representations between positive samples, and minimize the similarity of style representations between negative samples. Assume that the number of texts in each training batch is N, the set of positive samples is J, and the set of negative samples is M. Define the loss function as follows: Where sim(·,·) represents the cosine similarity function and τ is the temperature coefficient, which is used to control the smoothness or sharpness of the distribution; Step B22: Introduce the style classification task and use cross-entropy loss to ensure that the extracted style representation can effectively distinguish different styles: Among them, C is a style classifier pre-trained on dataset D, s i is the real style category; Step B23: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using backpropagation to minimize the loss function to train the style information extractor.
6. The unsupervised text style transfer method based on style decoupling according to claim 5, characterized in that The specific implementation method of step B3 is: Enter the text t a The intermediate text generator is passed in. The intermediate text generator adopts a text generation model and inputs the text t a Go through two processes: encoding and decoding; Enter text a First, it is converted into a word vector representation. Each meaningful word is converted into a token and then input into the encoder of the text generation model for encoding: H=Encoder(t a ) in, For input text t a The hidden state matrix, L is the input text t a The length of , d is the hidden layer dimension of the generative model, that is, the feature dimension; In the decoding phase, a continuous decoding algorithm is used to generate continuous embedding vectors. At each time step, the decoder calculates the probability distribution of the current token and performs a weighted summation with the word embedding matrix: h t =Decoder(v t-1 ,H) v t =p t ·W Where W is the word embedding matrix with a size of d×K, where K represents the size of the vocabulary; h tt is the output of the decoder of the text generation model at time step t, where the value range of t is [1, L], and h i Indicates h t The i-th element of , when t = 1, v0 is the starting symbol vector; p t is the word probability distribution calculated by the softmax function; The generated embedding vector v t As input for the next time step, we get the new output: h t+1 =Decoder(v t ,H) After L time steps, we get P = {p1, p2, ..., p L }, and then generate the intermediate text t n : t n =argmax(P)·W Among them, the argmax function is used to obtain the word probability distribution p t The index of the word with the highest probability.
7. The unsupervised text style transfer method based on style decoupling according to claim 6, characterized in that The step B4 specifically includes the following steps: Step B41: Train the intermediate text generator through the style removal task; s n =C(t n ) in, is the intermediate loss, s n Is the generated text t n The style score of , which takes values of -1 (indicating negative), 1 (indicating positive) or 0 (indicating neutral), and the training goal is to make the style score as close to 0 as possible; Step B42: Calculate the original text t a With the middle text t n The cosine similarity between: in, is the semantic loss, h a and h n The original text t a and the middle text t n The semantic vector obtained by word embedding is trained to minimize the loss and ensure the semantic consistency between the two. Step B43: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using backpropagation to minimize the loss function to train the intermediate text generator.
8. The unsupervised text style transfer method based on style decoupling according to claim 7 is characterized in that: The step B5 specifically includes the following steps: Step B51: Use the style transfer module to transfer the intermediate text t n Style representation z with target style b To process; first, the intermediate text t n Processed by the encoder of the style transfer module to obtain the semantic representation of the text: H n =Encoder(t n ,z b ) in, For the middle text t n The hidden state matrix, L is the input text t n The length of , d is the hidden layer dimension of the generative model, that is, the feature dimension; x t =Decoder(x t-1 ,H n ) Among them, x t-1 Represents the hidden state obtained in the previous time step. When t=1, the hidden state x0 is the starting symbol vector; y t is the predicted probability distribution of the tth word calculated by the Softmax function at the current time step t; After L time steps, we get Y = {y1,y2,…,y L }, Jiner generates text after style transfer Step B52: adopting a reconstruction training strategy during the training of the style transfer module; The original text t a Convert to target style text Its losses for: Among them, G is the style transfer module, ||·|| represents the text similarity measurement, and the goal is to make With the original text a As close as possible to ensure semantic integrity; At the same time, to ensure that the generated text conforms to the target style, its loss for: Among them, s b The target style of the generated sample; Step B53: The calculated loss as well as The learning rate is updated through the gradient optimization algorithm, and the model parameters are iteratively updated using back propagation to minimize the loss function to train the model.
9. An unsupervised text style transfer system based on style decoupling, characterized by: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 8 can be implemented.