Text processing method, device and system and computer readable storage medium
Through text processing methods based on classification model and optimization model, negative comment texts are converted tone, which solves the problems of wasting information and affecting user mood in the prior art, and achieves the effect of reducing negative impact while retaining the value of information.
Patent Information
- Application Number
- CN202311577606.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-27
AI Technical Summary
When processing negative comment texts, the prior art usually adopts the method of deleting or replacing sensitive information, which leads to wasting information or affecting the user's mood, and cannot effectively reduce the negative impact of comments with bad tone.
Text processing methods based on classification model and optimization model are adopted. After obtaining the pending text, the classification model is used to pre-process the text to determine text that contains sensitive words but has value and negative text that does not contain sensitive words. Then tone conversion is performed through the optimization model to generate the target text.
On the premise of retaining valuable information, it can effectively reduce the negative impact of comment texts with bad tone, improve user experience, and reduce the negative social impact of online violence.
Smart Images

Figure CN120045713A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a text processing method, apparatus, system, and computer-readable storage medium. Background Art
[0002] In the Internet era, people often interact with netizens and express their opinions through text comments. Since the communication and dialogue among Internet users do not require face-to-face interaction and are more direct and fast, the probability that each person issues a negative comment on the content they oppose has increased significantly. Everyone may receive negative comments, and many Internet users are often troubled by sudden unfriendly or even sensitive comments, which may even cause psychological problems. The accumulation of negative comments can constitute cyber violence and cause bad social impacts.
[0003] Existing methods for processing negative comment texts usually directly delete the comment texts containing sensitive information or replace the sensitive information in the comment texts with special symbols. Directly deleting the comment texts containing sensitive information results in a small number of viewable comment texts and a waste of valuable information. Replacing the sensitive information in the comment texts with special symbols still affects the user's mood and only treats the symptoms, not the root cause. Therefore, how to convert the tone of the text while retaining the valuable information in the text to reduce the impact of comment texts with a bad tone is an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide a text processing method, apparatus, system, and computer-readable storage medium, which can convert the tone of the text while retaining the valuable information in the text to reduce the impact of comment texts with a bad tone.
[0005] In a first aspect, embodiments of this application provide a text processing method, which includes:
[0006] Obtain a text to be processed;
[0007] Classify the text to be processed based on a classification model to obtain a preprocessed text;
[0008] Optimize the preprocessed text based on an optimization model to obtain a target text.
[0009] In a second aspect, embodiments of this application provide a text processing apparatus, including:
[0010] An obtaining unit, configured to obtain a text to be processed;
[0011] A classification unit, configured to classify the text to be processed based on a classification model to obtain a preprocessed text;
[0012] An optimization unit for optimizing the preprocessed text based on an optimization model to obtain a target text.
[0013] In a third aspect, an embodiment of the present application further provides a text processing system, including a memory storing multiple instructions; a processor loads instructions from the memory to execute the steps of any text processing method provided by the embodiments of the present application.
[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps of any text processing method provided by the embodiments of the present application.
[0015] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps in any text processing method provided by the embodiments of the present application are implemented.
[0016] Adopting the solution of the embodiments of the application, obtain a text to be processed; classify the text to be processed based on a classification model to obtain a preprocessed text; optimize the preprocessed text based on an optimization model to obtain a target text. Determine the preprocessed text in the text to be processed through the classification model, and then through the optimization model, optimize the text that contains sensitive words but has value and the negative text that does not contain sensitive words in the preprocessed text to obtain a target text, so as to convert the tone of the text while retaining valuable information and reduce the impact of texts with a bad tone. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is a schematic flowchart of an embodiment of the text processing method provided in the embodiments of the present application;
[0019] Figure 2 is a schematic structural diagram of the classification model provided in the embodiments of the present application;
[0020] Figure 3 is a schematic flowchart of determining negative texts, pure sensitive word texts, and non-pure sensitive word texts provided in the embodiments of the present application;
[0021] Figure 4 is a schematic flowchart of determining negative texts and texts containing sensitive words provided in the embodiments of the present application;
[0022] Figure 5 is a schematic flowchart for determining non-pure sensitive word texts and pure sensitive word texts provided in an embodiment of the present application;
[0023] Figure 6 is a schematic structural diagram of a conversion model provided in an embodiment of the present application;
[0024] Figure 7 is a schematic structural diagram of a text processing device provided in an embodiment of the present application;
[0025] Figure 8 is a schematic structural diagram of a text processing system provided in an embodiment of the present application. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first", "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0027] An embodiment of the present application provides a text processing method, device, system and computer-readable storage medium.
[0028] Specifically, the embodiment of the present application will be described from the perspective of a text processing system, which can be specifically integrated in a text processing device, that is, the text processing method in the embodiment of the present application can be executed by the text processing system.
[0029] The text processing method provided in the embodiment of the present application can be applied to a text processing device, and the text processing device can include a background server corresponding to various application programs with text functions.
[0030] The following will be described in detail with reference to the accompanying drawings respectively. In the embodiment of the present application, the execution subject is taken as the text processing system. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown in the accompanying drawings.
[0031] Please refer to Figure 1, a specific process of this text processing method includes the following steps:
[0032] Step 101, obtain the text to be processed;
[0033] Step 102, classify the text to be processed based on a classification model to obtain a preprocessed text;
[0034] Step 103, optimize the preprocessed text based on an optimization model to obtain a target text.
[0035] In this embodiment, the text processing system obtains the text to be processed, classifies the text to be processed based on a pre-created classification model to determine the preprocessed text in the text to be processed; for the preprocessed text in the text to be processed, the text processing system optimizes the preprocessed text based on an optimization model to obtain a target text, where the optimization model includes a screening model and a conversion model. The text processing system inputs the preprocessed text into the screening model to determine the negative text with unfriendly tone and the text containing sensitive words in the preprocessed text; the text processing system removes the sensitive words in the non-pure sensitive word text in the text containing sensitive words; the text processing system performs tone conversion on the non-pure sensitive word text after removing sensitive words and the negative text with unfriendly tone based on a pre-created conversion model, so that the non-pure sensitive word text after removing sensitive words and the negative text with unfriendly tone are not overly friendly while being more friendly than the original text, and are beautified on the basis of retaining authenticity, and the target text is obtained and output. It should be noted that sensitive words may include: abusive words, indecent words, and malicious words modified by changing tones, disassembling characters, pinyin substitution, etc.;
[0036] The text processing system of this embodiment obtains the text to be processed, classifies the text to be processed based on a pre-created classification model to obtain a preprocessed text; inputs the preprocessed text into a pre-created screening model to determine the target preprocessed text in the preprocessed text; removes the sensitive words in the target preprocessed text to obtain a preprocessed text; performs tone conversion on the preprocessed text based on a pre-created conversion model to obtain a target text. Through the classification model and the screening model, the text containing sensitive words but with value and the negative text without sensitive words are screened out, and then the text is subjected to tone conversion through the conversion model to obtain a target text, so that the tone of the text is converted on the premise of retaining valuable information, and the impact brought by the text with a bad tone is reduced.
[0037] Specifically, each step is described in detail as follows:
[0038] Step 101, obtain the text to be processed;
[0039] In this step, the text processing system can obtain the text to be processed according to the input of relevant personnel or in an automated detection manner. Exemplarily, when a user comments on certain content, the text processing system automatically detects and obtains the comment text initiated by the user as the text to be processed.
[0040] Step 102: Classify the text to be processed based on a classification model to obtain preprocessed text;
[0041] In this step, after the text processing system obtains the text to be processed, it classifies the text to be processed based on a pre-created classification model to determine the preprocessed text in the text to be processed. It can be understood that generally, the text to be processed obtained by the text processing system contains multiple texts. The text processing system inputs each text into the classification model in turn to determine the preprocessed text in the text to be processed.
[0042] In a feasible example, the pre-created classification model is used to perform sentiment analysis on the comment content sent by the user, and the comment text is divided into three categories: positive text, neutral text, and negative text. Positive text refers to text with obvious feelings of liking, neutral text refers to intermediate stance text without emotional color, and preprocessed text refers to text with a malicious tone, containing words with a sense of disgust, sarcasm, or sensitive words such as abuse and personal attacks. For positive text and neutral text, the text processing system directly displays them in the text area. For negative text determined as the preprocessed text in the text to be processed, the text processing system needs to further process it.
[0043] Specifically, step 102 includes:
[0044] Step 1021: Perform word segmentation on the text to be processed to obtain a word segmentation set;
[0045] In this step, the text processing system performs word segmentation on the text to be processed to obtain a word segmentation set; it should be noted that word segmentation is to divide a continuous text into meaningful units, usually words or subwords. In natural language processing, word segmentation is an important preprocessing step; the methods of word segmentation processing include but are not limited to: 1. Dictionary-based method: Use a predefined dictionary for matching and split the text according to the words in the dictionary; 2. Rule-based method: Use rules and heuristic algorithms to perform segmentation based on the relationship between characters, such as maximum forward matching, maximum backward matching, etc.; 3. Statistics-based method: Use statistical models such as Hidden Markov Model (HMM) and Conditional Random Field (CRF) to learn the word boundaries in the sentence and perform word segmentation. The specific word segmentation method can be set according to the actual situation and will not be limited here.
[0046] Step 1022: Input the word segmentation set into the classification model to obtain the first text probability and the second text probability corresponding to the text to be processed;
[0047] In this step, the text processing system inputs the word segmentation set into the pre-created classification model to obtain the first text probability and the second text probability corresponding to the text to be processed. It can be understood that each text to be processed corresponds to a word segmentation set, and multiple texts to be processed correspond to multiple word segmentation sets. The text processing system inputs the word segmentation set corresponding to each text to be processed into the classification model in sequence to obtain the first text probability and the second text probability corresponding to the text to be processed. Among them, the first text probability is the non-negative text probability, and the second text probability is the negative text probability.
[0048] In a feasible example, the classification model is an ALBERT-CMNN language model, and the ALBERT-CMNN language model combines two technologies: ALBERT (A Lite BERT) and CMNN (Circular Multi-level Neural Network).
[0049] The training set of the classification model includes two parts: text and the original content targeted by the text. The combination of the two enables more specific analysis of specific problems in text sentiment classification. It can be a publicly available text dataset or manually annotated. Text data with sentiment polarity should be labeled, and then each sample is preprocessed, including denoising, word segmentation, etc. Note that when removing stop words, tone words that have an important impact on sentiment classification cannot be removed.
[0050] As Figure 2 shown, the classification model includes: an input layer, an ALBERT layer, a CMNN layer, a Softmax layer, and an output layer. The training process of the classification model is as follows: First, perform word segmentation on the text and the original content targeted by the text. Figure 2Let X(N) denote the Nth word segment in the text. Then, the text and the text data corresponding to the content are input into the ALBERT model through the input layer. The first semantic feature of the text is extracted by the multi-layer bidirectional Transformer encoder. The first semantic feature trained in this way makes full use of the relationship between the text and the original content and the context information of each sentence. According to the information provided by the original content of the text, it can better analyze the sentiment polarity of the text and distinguish the different meanings of the same word in different contexts. Secondly, the first semantic feature is input into the CNN layer (Convolutional Neural Network layer) in the CMNN layer to further extract the local features of the text. N convolutional kernels of different sizes are used to train the text features, and the second semantic feature of the text is obtained. Furthermore, the second semantic feature is input into the BiLSTM layer (Bidirectional Long Short-Term Memory Network layer) in the CMNN layer, and the forward and backward sub-networks are used for modeling respectively. Finally, the third semantic feature of the text is obtained. The third semantic feature is fully connected twice through the fully connected layer. The output dimension of the fully connected layer is three categories: positive, neutral, and negative (or two categories: non-negative and negative). Finally, the Softmax function is used to normalize the output of the fully connected layer to obtain the positive text probability, neutral text probability, and negative text probability (or non-negative probability and negative text probability) corresponding to the text, and then determine whether the text belongs to positive, neutral, or negative (or non-negative or negative).
[0051] After the classification model is trained, it is applied to the text processing system. The text processing system inputs the word segment set into the pre-created classification model to obtain the positive text probability, neutral text probability, and negative text probability (or the non-negative text probability and negative text probability) corresponding to the text to be processed.
[0052] Step 1023: Set the text to be processed with the second text probability greater than the first text probability as the preprocessed text.
[0053] In this step, the text processing system sets the text to be processed with the second text probability greater than the first text probability as the preprocessed text through the classification model. In a feasible example, the output dimension of the fully connected layer of the classification model is three categories: positive, neutral, and negative. At this time, the positive text probability, neutral text probability, and negative text probability corresponding to the text to be processed are obtained. The text processing system determines through the classification model that the negative text probability is greater than the positive text probability and the neutral text probability, and then determines that the negative text in the text to be processed is the preprocessed text. In another feasible example, the output dimension of the fully connected layer of the classification model is two categories: non-negative and negative. At this time, the first text probability and the second text probability corresponding to the text to be processed are obtained. The text processing system determines through the classification model that the second text probability is greater than the first text probability, and then determines that the negative text in the text to be processed is the preprocessed text.
[0054] Step 103: Optimize the preprocessed text based on the optimization model to obtain the target text.
[0055] In this step, after obtaining the preprocessed text, the text processing system optimizes the preprocessed text based on the optimization model to obtain the target text. Specifically, the optimization model includes a screening model and a conversion model. The text processing system inputs the preprocessed text into the screening model to determine the negative text and non-pure sensitive word text in the preprocessed text, and then inputs the negative text and non-pure sensitive word text into the conversion model to obtain the target text.
[0056] It should be noted that the preprocessed text, that is, the negative text, includes negative text, non-pure sensitive word text, and pure sensitive word text. Among them, (1) Negative text does not contain sensitive words, but the expression is uncomfortable, which is called negative text. This type of text only has a bad tone, but the views expressed are meaningful and have retention value. (2) Non-pure sensitive word text consists of abusive information and clear expressions. This type of text may be a long sentence. Although the abusive information contains sensitive words, the clear expressions have retention value. (3) Pure sensitive word text may be very short, only containing simple appellations and sensitive words. The content of this type of text itself does not contain valuable information and should be directly blocked. Since this solution needs to convert the tone of the text while retaining valuable information and reduce the impact of negative text, therefore, pure sensitive word text is directly deleted, and negative text and non-pure sensitive word text are further processed subsequently to obtain the target text.
[0057] Specifically, Step 103 includes:
[0058] Step 1031: Input the preprocessed text into the screening model, perform sensitive word detection on the preprocessed text through the screening model, and divide the preprocessed text into negative text and sensitive word-containing text;
[0059] Step 1032: Determine the sensitive word ratio of the sensitive word-containing text through the screening model, and determine the sensitive word-containing text with a sensitive word ratio not greater than the preset ratio threshold as non-pure sensitive word text.
[0060] In steps 1031 to 1032, after determining the preprocessed text, the text processing system inputs the preprocessed text into a pre-created screening model. The sensitive word detection module in the screening model detects sensitive words in the preprocessed text, and divides the preprocessed text into negative text and text containing sensitive words. Then, the text processing system determines the proportion of sensitive words in the text containing sensitive words through the text containing sensitive word classification module in the screening model, and determines the text containing sensitive words with a sensitive word proportion not greater than the preset proportion threshold as non-pure sensitive word text. It can be understood that detecting and classifying in two steps can ensure the highest purity after classification, avoid misclassification of text near the critical threshold, and improve the accuracy of dividing negative text and non-pure sensitive word text.
[0061] In a feasible example, as Figure 3 shown, CFM is the screening model. CFM includes DACM and CACM, where DACM is the sensitive word detection module and CACM is the text containing sensitive word classification module. The above modules are all pre-trained. The text processing system first detects sensitive words in the preprocessed text through DACM, and divides the preprocessed text into negative text and text containing sensitive words. Then, through CACM, the text containing sensitive words is divided into pure sensitive word text and non-pure sensitive word text.
[0062] Specifically, step 1031 includes:
[0063] Step 10311, detecting sensitive words in the preprocessed text through the screening model;
[0064] Step 10312, if the sensitive word detection result is that the preprocessed text contains sensitive words, then determine the text containing sensitive words as the text containing sensitive words;
[0065] Step 10313, if the sensitive word detection result is that the preprocessed text does not contain sensitive words, then perform suspected sensitive word detection and context sensitive word detection on the preprocessed text through the screening model, and divide the preprocessed text into the negative text and the text containing sensitive words according to the suspected sensitive word detection result and the context sensitive word detection result.
[0066] In steps 10311 to 10313, the text processing system detects sensitive words in the preprocessed text through the sensitive word detection module in the screening model. If the sensitive word detection result indicates that the preprocessed text contains sensitive words, the text containing sensitive words is determined as the text with sensitive words; if the sensitive word detection result indicates that the preprocessed text does not contain sensitive words, the screening model further performs suspected sensitive word detection and context-sensitive word detection on the preprocessed text, and divides the preprocessed text into negative text and text with sensitive words according to the suspected sensitive word detection result and the context-sensitive word detection result. It can be understood that after determining that the preprocessed text does not contain sensitive words, optionally, the text processing system simultaneously performs suspected sensitive word detection and context-sensitive word detection on the preprocessed text through the screening model, and then divides the preprocessed text into negative text and text with sensitive words according to the suspected sensitive word detection result and the context-sensitive word detection result; optionally, the text processing system first performs suspected sensitive word detection on the preprocessed text through the text screening module to obtain the suspected sensitive word detection result, and then determines whether context-sensitive word detection is required according to the suspected sensitive word detection result. After determining that context-sensitive word detection is required, context-sensitive word detection is performed to obtain the context-sensitive word detection result, and then the preprocessed text is divided into negative text and text with sensitive words.
[0067] It should be noted that sensitive word detection is simply to detect each word in the preprocessed text to determine whether the preprocessed text contains sensitive words; suspected sensitive word detection is to detect malicious sensitive words modified by changing tones, splitting characters, pinyin substitution, etc.; context-sensitive word detection is to detect each word in the preprocessed text according to the semantic features of the context of the entire preprocessed text to determine whether the preprocessed text contains sensitive words.
[0068] Furthermore, step 10313 includes:
[0069] Step 103131, perform suspected sensitive word detection on the preprocessed text through the screening model to obtain the suspected sensitive word detection result, and the suspected sensitive word detection result includes the first detection result and the second detection result;
[0070] Step 103132, if both the first detection result and the second detection result indicate that the preprocessed text does not contain sensitive words, determine the text without sensitive words as negative text;
[0071] Step 103133, if both the first detection result and the second detection result indicate that the preprocessed text contains sensitive words, determine the text containing sensitive words as the text with sensitive words;
[0072] Step 103134, if the first detection result or the second detection result indicates that the preprocessed text contains a sensitive word, then perform context-sensitive word detection on the preprocessed text through the screening model;
[0073] Step 103135, if the context-sensitive word detection result indicates that the preprocessed text contains a sensitive word, then determine the text containing the sensitive word as the text with sensitive words;
[0074] Step 103136, if the context-sensitive word detection result indicates that the preprocessed text does not contain a sensitive word, then determine the text without the sensitive word as the negative text.
[0075] In steps 103131 to 103136, as Figure 4 shown, the text processing system performs sensitive word detection on the preprocessed text. If the sensitive word detection result indicates that the preprocessed text contains a sensitive word, then determine the text containing the sensitive word as the text with sensitive words; if the sensitive word detection result indicates that the preprocessed text does not contain a sensitive word, then perform suspected sensitive word detection on the preprocessed text through the screening model to obtain the suspected sensitive word detection result, and the suspected sensitive word detection result includes the first detection result (f1) and the second detection result (f2); if both the first detection result and the second detection result indicate that the preprocessed text does not contain a sensitive word (Not(f1 and f2)), then determine the text without the sensitive word as the negative text; if both the first detection result and the second detection result indicate that the preprocessed text contains a sensitive word (f1 and f2), then determine the text containing the sensitive word as the text with sensitive words; if the first detection result or the second detection result indicates that the preprocessed text contains a sensitive word (f1 or f2), then perform context-sensitive word detection on the preprocessed text through the screening model; if the context-sensitive word detection result indicates that the preprocessed text contains a sensitive word, then determine the text containing the sensitive word as the text with sensitive words; if the context-sensitive word detection result indicates that the preprocessed text does not contain a sensitive word, then determine the text without the sensitive word as the negative text.
[0076] In a feasible example, the suspected sensitive word detection can be implemented through the FastText model. FastText is a shallow neural network model based on the bag-of-words model. It regards the input text as a set of words, and first averages the word vectors and then performs classification; the context-sensitive word detection can be implemented through the BERT model. The BERT model is a deep bidirectional encoder based on the Transformer architecture and can consider context information simultaneously.
[0077] Specifically, step 1032 includes:
[0078] Step 10321: Determine the number of sensitive words and the total number of words in the text containing sensitive words through the screening model, and calculate the sensitive word ratio of the text containing sensitive words based on the number of sensitive words and the total number of words.
[0079] Step 10322: Compare the sensitive word ratio with a preset ratio threshold, and determine the text containing sensitive words with a sensitive word ratio not greater than the preset ratio threshold as the non-pure sensitive word text.
[0080] In Steps 10321 to 10322, after the text processing system determines the text containing sensitive words in the preprocessed text, usually the text containing sensitive words will include multiple texts. The text processing system processes each text containing sensitive words through the screening model in turn, determines the number of sensitive words and the total number of words in each text containing sensitive words, calculates the sensitive word ratio of each text containing sensitive words based on the number of sensitive words and the total number of words, and then compares the sensitive word ratio of each text containing sensitive words with the preset ratio threshold. The text containing sensitive words with a sensitive word ratio not greater than the preset ratio threshold is determined as the non-pure sensitive word text, and the text containing sensitive words with a sensitive word ratio greater than the preset ratio threshold is determined as the pure sensitive word text.
[0081] In a feasible example, as Figure 5 shown, the Llama-13B model is used to classify the text containing sensitive words. The Llama-13B model identifies and calculates the number of sensitive words and the total number of words in each text containing sensitive words, calculates the sensitive word ratio of each text containing sensitive words based on the number of sensitive words and the total number of words, and then divides the text containing sensitive words into non-pure sensitive word text and pure sensitive word text. The text processing system discards the pure sensitive word text and continues to process the non-pure sensitive word text in the next step.
[0082] It should be noted that the Llama-13B model is trained based on the Llama pre-trained model, and the fine-tuning task is defined as dividing sentences containing sensitive words into two categories with different sensitivity levels according to the sentiment polarity. The training process has two steps:
[0083] ① Dataset: The publicly available Chinese text corpus can be used, and sensitive words are added to the corpus in the dataset for processing. Suppose the total number of words in the processed corpus is N, and the number of sensitive words is n. Define the sensitive word ratio as n / N. The higher the sensitive word ratio of a sentence, the greater the probability of being predicted as a pure sensitive word text, and vice versa.
[0084] ② Model training: Convert the dataset into the format required by Llama, and load a pre-trained language model such as Llama-13B. Due to the complexity of Chinese sentence patterns and grammar, a larger model, more fine-tuning data, a longer fine-tuning time, and a higher learning rate may be required. Thus, a model that can distinguish between non-pure sensitive word texts and pure sensitive word texts is obtained.
[0085] Step 1033, perform tone conversion on the negative text and the non-pure sensitive word text with sensitive words removed through the conversion model to obtain the target text.
[0086] In this step, the text processing system identifies and removes the corresponding sensitive words in the non-pure sensitive word text to obtain the non-pure sensitive word text with sensitive words removed. The text processing system performs tone conversion on the negative text and the non-pure sensitive word text with sensitive words removed based on a pre-created conversion model to obtain the target text. Among them, the conversion model can be an improved Cycle-GAN model. The conversion model performs text style transfer on the negative text and the non-pure sensitive word text with sensitive words removed, so that the tone of the negative text and the non-pure sensitive word text with sensitive words removed can become relatively friendly and reduce aggression, while ensuring that the tone conversion is not overly friendly to prevent a significant contradiction with the original text's emotional style.
[0087] Further, before step 103, it includes:
[0088] Step a, obtain the first training text and the second training text, and create a pre-trained conversion model;
[0089] In this step, the text processing system obtains the first training text and the second training text, and creates a pre-trained conversion model. Among them, the first training text includes negative text and non-pure sensitive word text with sensitive words removed, the second training text is the text after the first training text is manually tone-converted, and the pre-trained conversion model is an improved Cycle-GAN model.
[0090] Step b, input the first training text into the first generator in the pre-trained conversion model to generate the first training conversion text, and input the first training conversion text into the second generator in the pre-trained conversion model to generate the second training conversion text;
[0091] In this step, as Figure 6 shown, the text processing system inputs the first training text (real_X) into the first generator (generator G) in the pre-trained conversion model to generate the first training conversion text (fake_Y), and inputs the first training conversion text (fake_Y) into the second generator (generator F) in the pre-trained conversion model to generate the second training conversion text (Cycle_X).
[0092] Step c: Input the second training text into the second generator to generate a third training conversion text, and then input the third training conversion text into the first generator to generate a fourth training conversion text.
[0093] In this step, as Figure 6 shown, the text processing system inputs the second training text (real_Y) into the second generator (Generator F) to generate a third training conversion text (fake_X), and then inputs the third training conversion text (fake_X) into the first generator (Generator G) to generate a fourth training conversion text (Cycle_Y).
[0094] Step d: Train the pre-trained conversion model based on the first training text, the second training text, the first training conversion text, the second training conversion text, the third training conversion text, and the fourth training conversion text to obtain a conversion model.
[0095] In this step, the text processing system trains the first generator, the second loss value, the third loss value of the similarity discriminator, and the scoring discriminator in the pre-trained conversion model based on the first training text, the second training text, the first training conversion text, the second training conversion text, the third training conversion text, and the fourth training conversion text to obtain a conversion model. Specifically, as Figure 6 shown, the first generator (Generator G), the second generator (Generator F), the similarity discriminator (Discriminator DX), and the scoring discriminator (Discriminator DY1 and Discriminator DY2).
[0096] Further, step d includes:
[0097] Step d1: Calculate the first loss value of the first generator based on the first training text and the second training conversion text.
[0098] Step d2: Calculate the second loss value of the second generator based on the second training text and the fourth training conversion text.
[0099] Step d3: Calculate the third loss value of the similarity discriminator in the pre-trained conversion model based on the first training text and the third training conversion text.
[0100] Step d4: Calculate the fourth loss value of the scoring discriminator in the pre-trained conversion model based on the second training text and the first training conversion text.
[0101] Step d5: Obtain a target loss value according to the first loss value, the second loss value, the third loss value, the fourth loss value, and a preset loss function.
[0102] Step d6: Based on the target loss value, perform backpropagation to update the parameters of the pre-trained conversion model until the target loss value meets the preset conditions, then the training is completed to obtain the conversion model.
[0103] In steps d1 to d6, the preset loss function of the pre-trained conversion model consists of adversarial Loss and cycle consistency Loss:
[0104]
[0105] where X is the first training text, Y is the second training text, G is the first loss value of the first generator, F is the second loss value of the second generator, and D X is the third loss value of the similarity discriminator, and D Y1 and D Y2 are the fourth loss values of the two scoring discriminators. λ is a training parameter that can be determined according to experimental details. To train the optimal generators G and F so that generator G can achieve the migration from X to Y as much as possible, and generator F can achieve the migration from Y to X as much as possible, and at the same time hope that the generations of the two generators can be reciprocal, the training objective of the preset loss function is:
[0106]
[0107] ①adversarial Loss: Discriminator cross-entropy loss. To make the data generated by the generator as close as possible to the real data distribution, set adversarial Loss, named L GAN , which consists of the Loss of three discriminators:
[0108]
[0109] For D Y1 :
[0110]
[0111] D Y2 Similarly. The improved Cycle GAN is the same as GAN. G is used to achieve X-Y. During training, try to make G(X) as close as possible to Y. The discriminators D Y1 , D Y2 discriminate the authenticity of samples, and the larger the Loss of D Y1 and D Y2 is, the more real the fake_Y is (in the present invention, it means that the converted tone reaches the relatively friendly standard). Therefore, the training objective of the model should be to make the Loss of G as small as possible, and the Loss of D Y1 , D Y2 as large as possible, that is Similarly for F.
[0112] ② Cycle consistency Loss: To ensure that the text style can be migrated back, a reconstruction loss is set to ensure that the samples generated by the two generators do not contradict each other. It is expected that during generation, x → G(x) → F(G(x)) ≈ x, which is called forward consistency. Similarly, it is expected that y → f(y) → G(F(y)) ≈ y, which is called backward consistency. To ensure consistency as much as possible, the cycle consistency Loss is set as:
[0113]
[0114] It should be noted that the Cycle-GAN model has a symmetric network structure, and each structure consists of two generators and a discriminator, which can achieve text style conversion. In this module, it is named as follows: The first generator G is used to convert the preprocessed text from an unfriendly tone X (referred to as real_X) to a relatively friendly neutral tone Y (Y generated by the generator G is called fake Y). The second generator F is used to convert the neutral tone text back to an unfriendly tone X (referred to as Cycle_X). The L2 Loss of minimizing the reconstruction loss is applied to the X-texts at both ends to ensure that the text content generated by F is consistent with the input text; the two discriminators D Y and D X are respectively used to judge the quality of the tone conversion effects of the two generators. Similarly for DX in the symmetric network structure.
[0115] The improved Cycle-GAN model is as Figure 6 shown. To control the emotional style of the target text after tone conversion within a certain range, two discriminators D Y can be set at the position of the discriminator D of the original Cycle-GAN model. Y1 and D Y2 . Both use the same training set, which consists of texts with tones ranging from good to bad. D Y1 should give a low score to texts with overly enthusiastic and friendly tones, and D Y2To give a low score to a text with an unfriendly tone, set the same threshold range for the two discriminators. If at least one discriminator gives a score outside this threshold range for the text, it means that the text after tone conversion is either too friendly or unfriendly. Then, it fails the discriminator and is regarded as an unqualified generation. Only when the scores given by both discriminators for the text after tone conversion are within the threshold range is it regarded as a qualified generation, and these two discriminators need to be co-trained. With this structure, the purpose of tone conversion can be achieved, and it can be ensured that the tone will not be overly friendly when it is more friendly than the original text, achieving beautification on the basis of retaining authenticity.
[0116] The text processing system of this embodiment obtains the text to be processed, classifies the text to be processed based on a pre-created classification model to determine the preprocessed text in the text to be processed; inputs the preprocessed text into a pre-created screening model to determine the target preprocessed text in the preprocessed text; removes sensitive words from the target preprocessed text to obtain the preprocessed text; performs tone conversion on the preprocessed text based on a pre-created conversion model to obtain the target text. Through the classification model and the screening model, texts that contain sensitive words but are valuable and negative texts that do not contain sensitive words are screened out. Then, through the conversion model, the text is tone-converted to obtain the target text, enabling the tone of the text to be converted while retaining valuable information and reducing the impact of texts with a bad tone.
[0117] This embodiment also provides a text processing device, which may specifically include a background server corresponding to various application programs with text functions, such as Figure 7 As shown, the text processing device may include:
[0118] An acquisition unit 1001, configured to acquire the text to be processed;
[0119] A classification unit 1002, configured to classify the text to be processed based on a classification model to obtain preprocessed text;
[0120] An optimization unit 1003, configured to optimize the preprocessed text based on an optimization model to obtain the target text.
[0121] In an optional example, the classification unit is further configured to:
[0122] Perform word segmentation on the text to be processed to obtain a word segmentation set;
[0123] Input the word segmentation set into the classification model to obtain a first text probability and a second text probability corresponding to the text to be processed;
[0124] Set the text to be processed with the second text probability greater than the first text probability as the preprocessed text.
[0125] In an optional example, the optimization unit is further configured to:
[0126] Input the preprocessed text into the screening model, perform sensitive word detection on the preprocessed text through the screening model, and divide the preprocessed text into negative text and text containing sensitive words;
[0127] Determine the sensitive word ratio of the text containing sensitive words through the screening model, and determine the text containing sensitive words with a sensitive word ratio not greater than a preset ratio threshold as non-pure sensitive word text;
[0128] Perform tone conversion on the negative text and the non-pure sensitive word text with sensitive words removed through the conversion model to obtain the target text.
[0129] In an optional example, the optimization unit is further configured to:
[0130] Perform sensitive word detection on the preprocessed text through the screening model;
[0131] If the sensitive word detection result is that the preprocessed text contains sensitive words, determine the text containing sensitive words as text containing sensitive words;
[0132] If the sensitive word detection result is that the preprocessed text does not contain sensitive words, perform suspected sensitive word detection and context sensitive word detection on the preprocessed text through the screening model, and divide the preprocessed text into the negative text and text containing sensitive words according to the suspected sensitive word detection result and the context sensitive word detection result.
[0133] In an optional example, the optimization unit is further configured to:
[0134] Perform suspected sensitive word detection on the preprocessed text through the screening model to obtain a suspected sensitive word detection result, where the suspected sensitive word detection result includes a first detection result and a second detection result;
[0135] If both the first detection result and the second detection result are that the preprocessed text does not contain sensitive words, determine the text not containing sensitive words as negative text;
[0136] If both the first detection result and the second detection result are that the preprocessed text contains sensitive words, determine the text containing sensitive words as text containing sensitive words;
[0137] If the first detection result or the second detection result is that the preprocessed text contains sensitive words, perform context sensitive word detection on the preprocessed text through the screening model;
[0138] If the context-sensitive word detection result indicates that the preprocessed text contains sensitive words, the text containing sensitive words is determined as the sensitive-word-containing text;
[0139] If the context-sensitive word detection result indicates that the preprocessed text does not contain sensitive words, the text without sensitive words is determined as the negative text.
[0140] In an optional example, the optimization unit is further configured to:
[0141] Determine the number of sensitive words and the total number of words in the sensitive-word-containing text through the screening model, and calculate the sensitive-word ratio of the sensitive-word-containing text based on the number of sensitive words and the total number of words;
[0142] Compare the sensitive-word ratio with a preset ratio threshold, and determine the sensitive-word-containing text with the sensitive-word ratio not greater than the preset ratio threshold as the non-pure-sensitive-word text.
[0143] In an optional example, the conversion unit is further configured to:
[0144] Obtain a first training text and a second training text, and create a pre-trained conversion model;
[0145] Input the first training text into the first generator in the pre-trained conversion model to generate a first training conversion text, and input the first training conversion text into the second generator in the pre-trained conversion model to generate a second training conversion text;
[0146] Input the second training text into the second generator to generate a third training conversion text, and input the third training conversion text into the first generator to generate a fourth training conversion text;
[0147] Train the pre-trained conversion model based on the first training text, the second training text, the first training conversion text, the second training conversion text, the third training conversion text, and the fourth training conversion text to obtain a conversion model.
[0148] In an optional example, the text processing device further includes a training unit, and the training unit is configured to:
[0149] Obtain a first training text and a second training text, and create a pre-trained conversion model;
[0150] Input the first training text into the first generator in the pre-trained conversion model to generate a first training conversion text, and input the first training conversion text into the second generator in the pre-trained conversion model to generate a second training conversion text;
[0151] Inputting the second training text into the second generator to generate a third training conversion text, and inputting the third training conversion text into the first generator to generate a fourth training conversion text;
[0152] Based on the first training text, the second training text, the first training conversion text, the second training conversion text, the third training conversion text and the fourth training conversion text, the pre-trained conversion model is trained to obtain a conversion model.
[0153] In an optional example, the training unit is further configured to:
[0154] Calculating a first loss value of the first generator based on the first training text and the second training conversion text;
[0155] Calculating a second loss value of the second generator based on the second training text and the fourth training conversion text;
[0156] Calculating a third loss value of a similarity discriminator in the pre-trained conversion model based on the first training text and the third training conversion text;
[0157] Calculating a fourth loss value of a scoring discriminator in the pre-trained conversion model based on the second training text and the first training conversion text;
[0158] Obtaining a target loss value according to the first loss value, the second loss value, the third loss value, the fourth loss value and a preset loss function;
[0159] Back propagation is performed based on the target loss value to update the parameters of the pre-trained conversion model until the target loss value meets the preset conditions, and the training is completed to obtain the conversion model.
[0160] According to the scheme of this embodiment, the text to be processed is obtained, and the text to be processed is classified based on the pre-created classification model to determine the pre-processed text in the text to be processed; the pre-processed text is input into the pre-created screening model to determine the target pre-processed text in the pre-processed text; the sensitive words in the target pre-processed text are removed to obtain the pre-processed text; the tone of the pre-processed text is converted based on the pre-created conversion model to obtain the target text. Through the classification model and the screening model, the text containing sensitive words but valuable and the negative text without sensitive words are screened out, and then the tone of the conversion model text is converted to obtain the target text, so that the tone of the text is converted while retaining valuable information, reducing the impact of the pre-processed text.
[0161] Accordingly, the present application also provides a text processing system, such as Figure 8 As shown,Figure 8 This is a schematic structural diagram of the text processing system provided by an embodiment of the present application. The text processing system 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. Among them, the processor 1101 is electrically connected to the memory 1102. Those skilled in the art can understand that the structural diagram of the text processing system shown in the figure does not constitute a limitation on the text processing system, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0162] The processor 1101 is the control center of the text processing system 1100, connecting various parts of the entire text processing system 1100 through various interfaces and lines. By running or loading software programs and / or units stored in the memory 1102, and by calling data stored in the memory 1102, it executes various functions of the text processing system 1100 and processes data, thereby monitoring the text processing system 1100 as a whole. The processor 1101 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0163] In the embodiment of the present application, the processor 1101 in the text processing system 1100 will load the instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to execute any one of the text processing methods provided by the embodiments of the present application.
[0164] Optionally, as Figure 8 shown, the text processing system 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art can understand that Figure 7 the structural diagram of the text processing system shown in [the figure] does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0165] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1101, and can receive and execute the commands sent by the processor 1101. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event. Subsequently, the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 1103 can also be used as part of the input unit 1106 to implement the input function.
[0166] The radio frequency circuit 1104 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other electronic devices through wireless communication, and transmit and receive signals with the network device or other electronic devices.
[0167] The audio circuit 1105 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105, converted into audio data, and then the audio data is output to the processor 1101 for processing, and then sent to another electronic device through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device.
[0168] The input unit 1106 can be used to receive input digital, character information or user feature information (such as fingerprint, iris, face information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0169] The power supply 1107 is used to supply power to each component of the text processing system 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1107 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0170] Although Figure 8 not shown in the figure, the text processing system 1100 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0171] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0172] Those of ordinary skill in the art can understand that all or part of the steps in the above methods of the embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by the processor.
[0173] Therefore, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored, and the computer programs can be loaded by the processor to execute any text processing method provided by the embodiments of the present application.
[0174] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0175] Since the computer program stored in the computer-readable storage medium can execute any one of the text processing methods provided in the embodiments of the present application, the beneficial effects achievable by any one of the text processing methods provided in the embodiments of the present application can be realized. For details, see the previous embodiments and will not be elaborated here.
[0176] According to one aspect of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the various optional implementation manners in the above embodiments.
[0177] In the above embodiments of the text processing device, computer-readable storage medium, text processing system, and computer program product, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects brought by the above-described text processing device, computer-readable storage medium, computer program product, text processing system and their corresponding units can refer to the description of the text processing method in the above embodiments, and will not be elaborated here specifically.
[0178] The above has introduced in detail a text processing method, device, system, computer-readable storage medium and computer program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A text processing method, characterized in that, the method includes: obtaining the text to be processed; classifying the text to be processed based on a classification model to obtain a preprocessed text; optimizing the preprocessed text based on an optimization model to obtain a target text.
2. The text processing method according to claim 1, characterized in that, the classifying the text to be processed based on a classification model to obtain a preprocessed text includes: performing word segmentation on the text to be processed to obtain a word segmentation set; inputting the word segmentation set into the classification model to obtain a first text probability and a second text probability corresponding to the text to be processed; setting the text to be processed with the second text probability greater than the first text probability as the preprocessed text.
3. The text processing method according to claim 1, characterized in that, the optimization model includes a screening model and a conversion model, and the optimizing the preprocessed text based on the optimization model to obtain a target text includes: inputting the preprocessed text into the screening model, detecting sensitive words in the preprocessed text through the screening model, and dividing the preprocessed text into negative texts and sensitive word-containing texts; determining the sensitive word ratio of the sensitive word-containing texts through the screening model, and determining the sensitive word-containing texts with the sensitive word ratio not greater than a preset ratio threshold as non-pure sensitive word texts; performing tone conversion on the negative texts and the non-pure sensitive word texts with sensitive words removed through the conversion model to obtain a target text.
4. The text processing method according to claim 3, characterized in that, the detecting sensitive words in the preprocessed text through the screening model and dividing the preprocessed text into the negative texts and sensitive word-containing texts includes: detecting sensitive words in the preprocessed text through the screening model; if the sensitive word detection result is that the preprocessed text contains sensitive words, determining the text containing sensitive words as the sensitive word-containing text; if the sensitive word detection result is that the preprocessed text does not contain sensitive words, performing suspected sensitive word detection and context-sensitive word detection on the preprocessed text through the screening model, and dividing the preprocessed text into the negative texts and sensitive word-containing texts according to the suspected sensitive word detection result and the context-sensitive word detection result.
5. The text processing method according to claim 4, characterized in that, the performing suspected sensitive word detection and context-sensitive word detection on the preprocessed text through the screening model and dividing the preprocessed text into the negative texts and sensitive word-containing texts according to the suspected sensitive word detection result and the context-sensitive word detection result includes: performing suspected sensitive word detection on the preprocessed text through the screening model to obtain a suspected sensitive word detection result, and the suspected sensitive word detection result includes a first detection result and a second detection result; if both the first detection result and the second detection result are that the preprocessed text does not contain sensitive words, determining the text not containing sensitive words as the negative text; If both the first detection result and the second detection result indicate that the preprocessed text contains sensitive words, then the text containing sensitive words is determined as the text with sensitive words; If either the first detection result or the second detection result indicates that the preprocessed text contains sensitive words, then the preprocessed text is subjected to context-sensitive word detection by the screening model; If the context-sensitive word detection result indicates that the preprocessed text contains sensitive words, then the text containing sensitive words is determined as the text with sensitive words; If the context-sensitive word detection result indicates that the preprocessed text does not contain sensitive words, then the text without sensitive words is determined as the negative text.
6. The text processing method according to claim 3, wherein, the step of determining the sensitive word ratio of the text with sensitive words by the screening model and determining the text with sensitive words whose sensitive word ratio is not greater than the preset ratio threshold as the non-pure sensitive word text includes: determining the number of sensitive words and the total number of words in the text with sensitive words by the screening model, and calculating the sensitive word ratio of the text with sensitive words based on the number of sensitive words and the total number of words; comparing the sensitive word ratio with the preset ratio threshold, and determining the text with sensitive words whose sensitive word ratio is not greater than the preset ratio threshold as the non-pure sensitive word text.
7. The text processing method according to claim 3, wherein, before the method of performing tone conversion on the negative text and the non-pure sensitive word text with sensitive words removed by the conversion model to obtain the target text, the method includes: obtaining a first training text and a second training text, and creating a pre-trained conversion model; inputting the first training text into the first generator in the pre-trained conversion model to generate a first training conversion text, and inputting the first training conversion text into the second generator in the pre-trained conversion model to generate a second training conversion text; inputting the second training text into the second generator to generate a third training conversion text, and inputting the third training conversion text into the first generator to generate a fourth training conversion text; training the pre-trained conversion model based on the first training text, the second training text, the first training conversion text, the second training conversion text, the third training conversion text, and the fourth training conversion text to obtain a conversion model.
8. The text processing method according to claim 7, wherein, the step of training the pre-trained conversion model based on the first training text, the second training text, the first training conversion text, the second training conversion text, the third training conversion text, and the fourth training conversion text to obtain a conversion model includes: calculating a first loss value of the first generator based on the first training text and the second training conversion text; calculating a second loss value of the second generator based on the second training text and the fourth training conversion text; Calculate a third loss value of a similarity discriminator in the pre-training transformation model based on the first training text and the third training transformation text; Calculate a fourth loss value of a scoring discriminator in the pre-training transformation model based on the second training text and the first training transformation text; Obtain a target loss value according to the first loss value, the second loss value, the third loss value, the fourth loss value, and a preset loss function; Perform backpropagation based on the target loss value to update parameters of the pre-training transformation model until the target loss value meets a preset condition, and then complete the training to obtain a transformation model.
9. A text processing device, characterized in that, the device includes: an acquisition unit configured to acquire a text to be processed; a classification unit configured to classify the text to be processed based on a classification model to obtain a preprocessed text; an optimization unit configured to optimize the preprocessed text based on an optimization model to obtain a target text.
10. A text processing system, characterized in that, it includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps of the text processing method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the text processing method according to any one of claims 1-8.