Text processing method, computer device and storage medium
The pre-trained text polishing model polishes sentences, modifiers and vocabulary are solved, and the problem that text correction technology in the existing technology cannot be accurately polished is achieved, achieving efficient improvement of text content.
Patent Information
- Application Number
- CN202111544933.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-12-16
AI Technical Summary
Existing text correction techniques cannot accurately determine whether the statement is in line with the application scenario, and cannot effectively polish the statements that express inaccurately, resulting in difficulty in improving the quality of the text content.
The pre-trained text polishing model is used to obtain the text to be processed and to polish the modifiers and vocabulary to be optimized in natural language sentences. The GPT and BART models are used to polish the sentence-level, modified word-level and vocabulary-level, and combine excellent expression sentence recognition and synonym diagram for optimization.
It achieves efficient and accurate embellishment of text content, improves the gracefulness of the sentence and the connection with the context, and improves the overall quality of the text.
Smart Images

Figure CN114298031B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of language processing technology, and in particular to a text processing method, computer equipment, and storage medium. Background Art
[0002] Text polishing refers to the process of modifying and reworking the textual content of a document to enhance its richness and literary quality. Since text polishing requires not only basic language skills like vocabulary and grammar, but also more advanced language skills like idioms and rhetoric, automated text polishing is quite challenging.
[0003] Currently, text quality verification often relies on text correction technology. However, existing text correction technology can only identify whether rhetorical devices are used in a text. It cannot accurately determine whether a sentence is consistent with the application scenario, nor can it accurately polish an inaccurate sentence. Therefore, there is an urgent need for a method that can efficiently and accurately improve the quality of text content. Summary of the Invention
[0004] The present application provides a text processing method, computer device and storage medium, which can polish text at the sentence, sentence modifier and / or sentence vocabulary level, and can efficiently and accurately improve the quality of text content.
[0005] In a first aspect, the present application provides a text processing method, the method comprising: obtaining a text to be processed, the text to be processed comprising at least one natural language sentence;
[0006] The text to be processed is input into a pre-trained text polishing model, and the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence are polished to obtain a target text.
[0007] In a second aspect, the present application further provides a computer device, comprising:
[0008] memory and processor;
[0009] The memory is used to store computer programs;
[0010] The processor is used to execute the computer program and implement the steps of the text processing method described in the first aspect when executing the computer program.
[0011] In a third aspect, the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps of the text processing method described in the first aspect above.
[0012] This application discloses a text processing method, computer device, and storage medium. The method first obtains a text to be processed, the text comprising at least one natural language sentence. The method then inputs the text to be processed into a pre-trained text polishing model, polishing the natural language sentence, its modifiers, and / or the vocabulary to be optimized within the sentence to obtain a target text. The method can polish text at the sentence, sentence modifier, and / or sentence vocabulary levels, effectively and accurately improving the quality of text content. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0014] Figure 1 This is a schematic diagram of an application scenario architecture of a text processing method disclosed in an embodiment of the present application;
[0015] Figure 2 is a structural diagram of a text processing system provided in an embodiment of the present application;
[0016] Figure 3 This is a schematic diagram of an application scenario of a text processing method provided by another embodiment of the present application;
[0017] Figure 4 This is a schematic diagram of the implementation flow of the text processing method provided in one embodiment of the present application;
[0018] Figure 5 is a flowchart of an implementation of a text processing method provided by another embodiment of the present application;
[0019] Figure 6 yes Figure 5 The specific implementation flow chart of S502;
[0020] Figure 7 This is a schematic diagram of the principle of the first target network branch;
[0021] Figure 8 This is a schematic diagram of the principle of the second target network branch;
[0022] Figure 9This is a schematic diagram of the principle of the third target network branch;
[0023] Figure 10 It is a schematic block diagram of the structure of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0026] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0027] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0028] Embodiments of the present application provide a text processing method, computer device, and storage medium. The text processing method provided in the embodiments of the present application can be used to input a text to be processed into a pre-trained text polishing model, and perform polishing on natural language sentences, modifiers in natural language sentences, and / or vocabulary to be optimized in natural language sentences in the text to be processed, thereby efficiently and accurately improving the quality of the text content and obtaining a high-quality target text.
[0029] It should be noted that in the field of natural language processing, article polishing generally refers to the process of modifying and processing the text content in a text to make the content fuller and more literary. Because the expression of an article requires the mastery of basic language skills such as vocabulary and grammar before being able to write sentences with correct words and complete syntactic structures. If people do not accumulate and practice enough in language expression, do not use more advanced vocabulary such as idioms, and do not pay attention to the use of rhetoric devices such as metaphors and personification, the sentences they write will have problems such as dry content and bland text. Therefore, language expression ability is a relatively important factor affecting the ability of writers.
[0030] Currently, to help improve students' writing skills, relatively mature essay correction and essay recommendation technologies have been developed. Essay correction technology can identify incorrect sentences in an essay, but it requires essay recommendation technology to search the resource library for sentences that are semantically similar to the current sentence. However, essay recommendation technology is limited by the size of the resource library and its use of contextual information. The sentences obtained through text matching technology are generally difficult to fit the context of the current essay, which to a certain extent affects students' learning and perception of these sentences. Therefore, automatic essay polishing technology has been proposed, hoping to combine automatic essay polishing technology with essay correction technology to rewrite some sentences so that the sentences in the entire article are consistent with the current context and are more beautiful and fluent than the original sentences.
[0031] In view of this, the present application proposes a text processing method, computer equipment and storage medium, which realizes comprehensive polishing of text content from three levels: sentences, fragments and words, so as to achieve the purpose of improving the quality of text content.
[0032] See also Figure 1 , Figure 1 This is a schematic diagram of the application scenario architecture of the text processing method disclosed in one embodiment of this application. Figure 1 As shown, the text processing method can be applied to a computer device 10, which has a text processing system 11 integrated therein. The computer device 10 can be a server or a terminal device. The server can be a remote server, a cloud server, or a server cluster, which can be used to run the text processing system 11. The terminal device can be a personal computer, a notebook, a PAD, a robot, or a handheld smart terminal device with sufficient computing power. The terminal device can also be used to run the text processing system 11. The text processing system 11 is an application program integrated on the computer device 10 with text processing capabilities.
[0033] It should be understood that the text processing method described in the embodiments of the present application can be applied to all application scenarios in which a text processing system 11 is integrated into a computer device 10 and text content is polished through the text processing system 11.
[0034] For example, see Figure 2 As shown, Figure 2 This is a structural diagram of the text processing system provided by the embodiment of the present application. Figure 2 It can be seen that the text processing system 11 includes:
[0035] An acquisition module 111 is configured to acquire a text to be processed, wherein the text to be processed includes at least one natural language sentence;
[0036] The polishing module 112 is configured to input the text to be processed into a pre-trained text polishing model, perform polishing on the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence, to obtain a target text.
[0037] Due to the functions of the above modules, the text polishing system provided in the embodiment of the present application can polish the text from the sentence, sentence modifiers and / or sentence vocabulary level, and can efficiently and accurately improve the quality of the text content.
[0038] Furthermore, the text processing system 11 may be composed of multiple subsystems. For example, the text processing system 11 includes a training subsystem 101 and a processing subsystem 102. It should be understood that the training subsystem 101 and the processing subsystem 102 may be integrated into a computer program. In the computer device 10, the text processing system 11 serves as an application program for polishing text content. Alternatively, in the computer device 10, the training subsystem 101 and the processing subsystem 102 may be integrated into the same application program for polishing text content.
[0039] It should be understood that the training subsystem 101 and the processing subsystem 102 can also be integrated into the computer device 10 as two different applications, with the computer device 10 implementing their respective functions by invoking the corresponding applications. For example, the computer device 10 trains the text polishing model by invoking the application corresponding to the training subsystem 101, and polishes the text to be processed based on the text polishing model by invoking the application corresponding to the processing subsystem 102.
[0040] It should be understood that when the computer device 10 has high computing capabilities, such as a server or server cluster, the corresponding training subsystem 101 and processing subsystem 102 can both be integrated into the computer device 10. When the computer device 10 is a terminal device with limited computing capabilities, such as a handheld smart terminal device, it is possible to consider integrating only the processing subsystem 102 into the terminal device, while deploying the training subsystem 101 in the cloud that is communicatively connected to the terminal device.
[0041] For example, Figure 3 As shown, Figure 3 This is a schematic diagram of an application scenario of a text processing method provided by another embodiment of the present application. Figure 3 In the figure, the text processing method is implemented by the cloud platform 20 and the computer device 10. It should be noted that the cloud platform 20 includes a cloud data center and a cloud service platform (not shown in the figure). The cloud data center includes a large number of basic resources owned by the cloud service provider, such as preset language models; the computing resources included in the cloud data center can be a large number of computer devices, such as servers or server clusters.
[0042] In this embodiment, a preset language model is deployed on the cloud service platform, and a training subsystem is deployed on the cloud data center. The training subsystem can train the preset language model on the cloud service platform to obtain a text polishing model, thereby ensuring that the computer device 10 polishes the text to be processed based on the text polishing model.
[0043] In some embodiments, the training subsystem is used to pre-train the preset language model in the cloud service platform based on the training samples to obtain a pre-trained text polishing model, and send the pre-trained text polishing model to the computer device 10 so that the computer device 10 can polish the text to be polished, such as sentence-level polishing, modifier-level polishing and word-level polishing of the composition, thereby improving some excellent expression sentences, modifiers or word usage examples that fit the context, so as to achieve the purpose of assisting writers, such as students' independent learning and improving students' writing ability.
[0044] It should be noted that in the embodiments of the present application, in order to improve the quality of the text content to be processed, excellent expression sentence generation tasks, modifier supplementation tasks and vocabulary upgrading tasks are defined. Among them, the excellent expression sentence generation task includes: given the first half of a sentence, automatically completing the second half of the sentence, so that the entire sentence is beautiful, fluent, and fits the current context of the sentence. The modifier supplementation task includes: given a sentence, automatically predicting the position where text needs to be inserted in the sentence, and generating a modifier fragment and / or modifier that conforms to the semantics of the current sentence, and supplementing the generated modifier fragment and / or modifier to the position where text needs to be inserted; the vocabulary upgrading task is actually a task of replacing vocabulary in the text, that is, given a sentence, automatically predicting the words that need to be replaced, and giving one or more words that can be replaced in the sentence, and replacing the corresponding words that need to be replaced.
[0045] For example, as shown in Table 1, Table 1 is an exemplary description of the generation of excellent expression sentences, the supplementation of modifiers, and the application of vocabulary upgrading.
[0046]
[0047] Table 1
[0048] As shown in Table 1, the original sentence: She has a small body with a little fat, and is a little short. After the excellent expression sentence generation task is used to complete the text, the obtained sentence is: She has a small body with a little fat, and is a little small in stature, and looks very cute. The original sentence: You will find that the apple is covered with small white dots. After the modifier supplementation task is used to supplement the modifiers, the obtained sentence is: You will find that the apple is covered with small white dots like sesame seeds. The original sentence: We must persevere in the face of difficulties. After the vocabulary upgrading task is used to replace the text, the obtained sentence is: We must persevere in the face of difficulties.
[0049] Since there is currently no established automatic essay polishing system, the embodiments of this application mainly draw on some natural language generation technologies based on pre-trained models to implement related functions of text polishing. In the field of natural language processing, there are two common natural language generation technologies based on pre-trained models: one is the language model-based natural language generation technology, represented by the unsupervised learning generation (Gererate Pre-Training Model, GPT) model; the other is the conditional language model-based natural language generation technology, represented by the conditional generation (Bidirectional and Auto-Regressive Transformers, BART) model.
[0050] The structure of the GPT model is a multi-layer stacked TransformerDecoder. The GPT model uses a language model task for pre-training, that is, the previous character in the text sequence is used to predict the next character, thereby achieving self-supervised training on large-scale text. The GPT model has been widely used in various natural language generation and natural language understanding tasks. The more representative GPT-based models include the DialoGPT model and the CPM model. DialoGPT is an application of the GPT model to the dialogue generation task, which has achieved competitive results on multiple dialogue generation task datasets. The CPM model is a Chinese version of the GPT model. This model uses large-scale multi-source Chinese text for pre-training and modifies the vocabulary of the original GPT model to make it more suitable for Chinese.
[0051] The BART model adopts a sequence-to-sequence model structure, using a multi-layer stack of Transformer Encoders as the encoder and a multi-layer stack of Tranformer Decoders as the decoder. To achieve self-supervised pre-training in the form of a conditional language model, the BART model first uses various methods of text corruption during pre-training to corrupt the original text sequence. The corrupted text sequence is then used as the input to the encoder and the uncorrupted text sequence as the output to the decoder. These methods of text corruption include randomly replacing characters in the text sequence with "[MASK]", randomly deleting characters in the text sequence, and scrambling the order of sentences in the text sequence. This conditional language model-based pre-training model is more suitable for sequence-to-sequence natural language generation tasks such as machine translation, grammatical error correction, and text summarization.
[0052] For the task of generating excellent expression sentences, directly using the GPT model may result in relatively smooth generation results. However, the beauty of these generation results and the degree of cohesion with the context are difficult to guarantee. Directly using the GPT model is likely to generate a sentence that is not as beautiful as the original sentence and is far from the original sentence in content, which certainly cannot achieve the purpose of text polishing. Therefore, on the one hand, in order to ensure the beauty of the generation results, the excellent expression sentence recognition model is first used to extract excellent expression sentences with high confidence from the massive text corpus as the corpus for subsequent fine-tuning of the GPT model. On the other hand, in order to ensure the degree of cohesion between the generation results and the context, a keyword information-guided fine-tuning method is used to fine-tune the GPT model, which, to a certain extent, ensures the degree of cohesion between the generation results and the context.
[0053] In addition, for the task of supplementing modified segments and modifiers, there are two technical problems that need to be solved: the first problem is how to determine the insertion position of the modified segments and modifiers; the second problem is how to ensure that the generated results are smooth and consistent with the semantics of the current text. At present, a pointer network is often used to predict a position, and then the vector representation of this position is extracted and input into a randomly initialized multi-layer stacked Transformer Decoder to generate the inserted segment. The main disadvantage of the above scheme is that the Transformer Decoder it uses is randomly initialized, and its training cost is high. Only by relying on large-scale data for training can better generation results be obtained. On the other hand, the method of superimposing the Transformer Decoder on the BERT model increases the complexity of the model and affects the training and reasoning speed of the model. In view of the above two problems, in the embodiment of the present application, a pointer generation network based on BART is used to complete the task of supplementing modified segments and modifiers. Since the encoder and decoder of BART are pre-trained on large-scale texts, only a small amount of data is needed to obtain good generation results.
[0054] For the vocabulary upgrading task, the first problem that needs to be solved is the problem of obtaining synonyms, that is, how to obtain synonyms of the words to be upgraded. The next problem that needs to be solved is the problem of sorting these synonyms, that is, how to select some words that are suitable for the current context and have better expression effects after use from these synonyms. For the problem of obtaining synonyms, a common technical solution is usually to use a method of calculating word vector similarity to obtain them, but directly calculating the word vector similarity between the word to be upgraded and all the words in the vocabulary will affect the calculation speed and may introduce some noise words. And for the task of sorting synonyms, when sorting synonyms by calculating the perplexity of the language model, the perplexity can only guarantee that the sentence will be more fluent after using the word, but cannot guarantee that the sentence will be more beautiful after using the word. In view of the above two problems, in an embodiment of the present application, synonyms are obtained through a synonym word graph constructed by word interpretation, which improves the efficiency and effect of synonym acquisition, and by using the GPT model and the excellent expression sentence recognition model to score and sort these synonyms, the expression quality of the text content can be improved in many aspects.
[0055] See also Figure 4 As shown, Figure 4 This is a schematic diagram of the implementation flow of the text processing method provided by an embodiment of the present application. The text processing method provided by this embodiment can be Figure 1 The computer device 10 shown is completed. The details are as follows:
[0056] S401: Obtain a text to be processed, where the text to be processed includes at least one natural language sentence.
[0057] Among them, the text to be processed includes but is not limited to compositions written by primary school students, articles by novice writers, articles by business writing enthusiasts, etc. In the text to be processed, due to the text author's inadequate understanding of sentence grammar, common vocabulary, allusions, modifiers, etc., there may be grammatical errors, misuse of modifiers, and incorrect vocabulary usage in the content of the text to be processed. In an embodiment of the present application, the above-mentioned problems existing in the text to be processed can be optimized through a pre-trained text polishing model to obtain a text with excellent expression sentences. Specifically, the text to be processed includes at least one natural language sentence, and a high-quality text can be obtained by processing the natural language sentence through a pre-trained text polishing model.
[0058] S402: Input the text to be processed into a pre-trained text polishing model, perform polishing on the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence, to obtain a target text.
[0059] The pre-trained text polishing model is obtained by training a preset language model based on a preset number of training samples. Specifically, the preset language model includes: a first network architecture, a second network architecture, and a third network architecture.
[0060] The first network architecture primarily produces sentence-level polishing results. It comprises a keyword-guided GPT sentence generation network model. This keyword-guided GPT generation model primarily focuses on generating descriptive sentences about scenery, appearance, and demeanor. It can polish natural language sentences at the sentence level, producing elegant sentences and improving the effectiveness of their expression.
[0061] The second network architecture is the Modified Segment and Modifier Supplement Generation Module. This module, composed of a BART-based pointer generation network model, primarily generates segment-level or word-level polishing results. This module focuses on inserting metaphorical modifiers, adjectives, adverbs, and other similar modifiers into sentences. This module can polish natural language sentences at the modifier level, introducing appropriate modifiers and achieving the effect of vivid expression through these modifiers.
[0062] The third network architecture is the vocabulary generation module. The main function of the vocabulary generation module is to produce word-level polishing results, and its component is a GPT vocabulary generation network model guided by replacement sentences. Among them, the GPT vocabulary generation network model guided by replacement sentences can include a synonym word graph and a synonym scoring network based on the BERT model and the excellent expression sentence recognition model. The synonym scoring network mainly focuses on upgrading verbs, adjectives, and adverbs. It can polish natural language sentences from the lexical level of the sentence, replace inappropriate words or common words in the text with more accurate and expressive words, so as to improve the quality of the text.
[0063] In some embodiments, the pre-trained text polishing model includes a first target network branch, a second target network branch, and a third target network branch that are independent of each other; the first target network branch is used to detect natural language sentences that need to be polished in the text to be processed, and polish the natural language sentences that need to be polished to obtain first target sentences; the second target network branch is used to detect target positions that need to be supplemented with modifiers in each of the natural language sentences, and supplement the target positions with modifiers to obtain second target sentences; the third target network branch is used to detect vocabulary to be optimized in each of the natural language sentences, and optimize the vocabulary to be optimized to obtain target vocabulary.
[0064] From the above analysis, it can be seen that the text processing method provided in the embodiment of the present application first obtains a text to be processed, wherein the text to be processed includes at least one natural language sentence; then the text to be processed is input into a pre-trained text polishing model, and the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence are polished to obtain a target text. The text can be polished at the sentence, sentence modifier, and / or sentence vocabulary levels, which can efficiently and accurately improve the quality of the text content.
[0065] See also Figure 5 As shown, Figure 5 This is a flowchart of the text processing method provided by another embodiment of the present application. Figure 5 It can be seen that this embodiment is Figure 4 Compared with the embodiment shown, the specific implementation process of S501 is the same as S401 and S503 is the same as S402. The difference is that S502 is included before S503. It should be understood that S502 and S501 are executed in parallel and can be executed one by one. The details are as follows:
[0066] S501: Acquire a text to be processed, where the text to be processed includes at least one natural language sentence.
[0067] S502: Obtain a preset number of training samples, train a preset language model, and obtain the text polishing model.
[0068] It should be understood that the preset language model includes a first network, a second network, and a third network having different polishing functions for the text. In the embodiments of the present application, the training processes of the first network, the second network, and the third network are exemplified respectively.
[0069] like Figure 6 As shown, Figure 6 yes Figure 5 The specific implementation flow chart of S502 is shown in FIG. Figure 6 It can be seen that S502 includes S5021 to S5023. The details are as follows:
[0070] S5021: Obtain a first preset number of first training samples, and train a preset first network architecture based on the first training samples to obtain the first target network branch.
[0071] The obtaining of a first preset number of first training samples includes: obtaining a first target number of text data; screening the text data based on a preset excellent text correction strategy to obtain the first preset number of excellent sentences, and using the excellent sentences as the first training samples.
[0072] Among them, the training of the first network architecture based on the first training sample to obtain the first target network branch includes: for any one of the excellent sentences, extracting keywords for the excellent sentence according to a preset keyword extraction algorithm to obtain at least one keyword; generating a first sequence of the excellent sentences and a second sequence of the keywords respectively; supplementing the second sequence to the first sequence to obtain a target sequence; inputting the target sequence corresponding to each of the excellent sentences into the first network architecture, training the first network architecture based on the target sequence, and obtaining the first target network branch.
[0073] It should be understood that the step of adding the second sequence to the first sequence to obtain the target sequence may include adding the second sequence before the first sequence, or adding the second sequence after the first sequence to obtain the target sequence.
[0074] In some embodiments, the training of the first network architecture based on the target sequence to obtain the first target network branch includes: training the first network architecture based on the target sequence and detecting the value of a first loss function; if the value of the first loss function is less than or equal to a first preset threshold, determining that the training of the first network architecture is completed and obtaining the first target network branch.
[0075] It should be understood that in order for the first target network branch to generate excellent expression sentences, it is first necessary to obtain training corpus of excellent expression sentences. In an embodiment of the present application, the training corpus of excellent expression sentences can be obtained through an excellent text correction strategy. Among them, the excellent text correction strategy uses pre-trained metaphor recognition models, personification recognition models, parallelism recognition models, quotation recognition models, description recognition models, excellent expression sentence recognition models, etc. to extract a preset number, for example, 3.5 million high-confidence excellent expression sentences as training corpus from massive text corpus. The training corpus of these excellent expression sentences can, to a certain extent, ensure that the first target network branch obtained by training can generate excellent expression sentences that meet the requirements.
[0076] Specifically, excellent text correction strategies bring together many pre-trained network models and automatically output sentences with rhetorical techniques.
[0077] For example, Figure 7 As shown, Figure 7 This is the principle diagram of the first target network branch. Figure 7 As can be seen, in this embodiment of the present application, the keyword in the sentence to be polished is "jujube tree, jujubes", and the sentence to be polished is: "The branches of the jujube tree are full of jujubes." After polishing through the first target network branch, the output excellent sentences include: "The branches of the jujube tree are full of bright red fruits"; "The branches of the jujube tree are full of big, red jujubes"; and "The branches of the jujube tree are full of heavy winter jujubes, and each jujube is as crystal clear as a ruby."
[0078] Specifically, the training process of the first target network branch includes: after obtaining the corpus of excellent expression sentences, a keyword information-guided method is used to fine-tune the GPT model.
[0079] Among them, for a sentence in the training set composed of excellent expression sentences, it is assumed that the sequence corresponding to the sentence is represented as: S = {w1, w2, ..., w n}, extract several keywords from the sentence through a keyword extraction algorithm, such as the graph-based Text Rank algorithm for text, and the keyword sequence is assumed to be: K = {k1, k2, ..., k m}, where m < n, n represents the number of words in the sentence S sequence, and m represents the number of keywords in the keyword sequence. Add the keyword sequence K to the front of the sentence sequence S, and use the "[SEP]" delimiter to separate the keyword sequence K and the sentence sequence S to obtain the final input sequence S of the fine-tuned GPT model * ={s1, s2, ..., s n+m+1}. The input sequence S corresponding to all statements *The result is sent to the GPT model for training, and the training of the GEP model is determined to be completed based on the first loss function of the GPT model.
[0080] For example, the first loss function can be expressed as:
[0081] h i =TransformerDecoder(s i )
[0082] y i+1 =Softmax(Linear(h i ))
[0083] loss i =CrossEntropy(y i+1 , s i+1 )
[0084] where h i Represents character s i The vector representation after encoding by GPT's encoding layer TransformerDecoder, y i+1 According to h i The next character predicted, loss i Represents the predicted character y i+1 and real characters i+1 The cross entropy loss value of .
[0085] It should be understood that after the fine-tuning of the GPT model is completed, it can be applied to the generation of excellent expression sentences, that is, the GPT model after fine-tuning is the first target network branch obtained by training.
[0086] It should be understood that when describing scenery, appearance, demeanor, etc., if excellent expressions are needed, the first target network branch can be used to identify descriptive sentences related to scenery, appearance, and demeanor in the corresponding text. Then, sentences with low confidence levels of excellent expressions among these descriptive sentences are selected as sentences requiring polishing. Finally, for each sentence requiring polishing, the keywords of this sentence are concatenated with the sentence using a delimiter to generate an excellent expression for that sentence.
[0087] It should be understood that in order to ensure the quality and diversity of the excellent expressions generated by the first target network branch, during the training process of the first target network branch, a keyword sequence reordering method can be used to generate multiple different keyword sequences, and the first target network branch can be trained to thereby obtain a first target network branch with a variety of polished results. Furthermore, during the training of the first target network branch, the perplexity of the GPT language model can be further used to sort the various polished results to select the optimal excellent expressions.
[0088] S5022: Obtain a second preset number of second training samples, and train a preset second network architecture based on the second training samples to obtain the second target network branch.
[0089] Among them, the obtaining of a second preset number of second training samples includes: obtaining a second target number of text data; screening the text data based on a preset excellent text correction strategy to obtain the second preset number of excellent sentences; extracting modifiers from the excellent sentences to obtain sentences without modifiers; generating a corpus pair based on the sentences without modifiers and the first position of the modifiers, and using the corpus pair as the second training sample.
[0090] In some embodiments, the second network architecture is trained based on the second training sample to obtain the second target network branch, including: for any pair of the corpus pairs, analyzing the sentence without modifiers based on the second network architecture to obtain the word vector of each word in the sentence without modifiers; predicting the second position of the modifier in the word vector based on a preset pointer network; calculating the cross entropy of the first position and the second position, and determining the insertion position of the modifier according to the cross entropy; generating a position vector according to the insertion positions of all the modifiers; inputting the position vector and the word vector into a preset encoder to predict the modifier at each insertion position; determining the end of training of the second network architecture according to the value of the second loss function to obtain the second target network branch.
[0091] In some embodiments, determining the end of training of the second network architecture based on the value of the second loss function to obtain the second target network branch includes: if the value of the second loss function is less than or equal to a preset threshold, determining the end of training of the second network architecture to obtain the second target network branch.
[0092] It should be understood that in order to train the second target network branch, it is first necessary to construct some sentences without modifiers and / or modifier fragments, and then to construct training corpus pairs from sentences without modifiers and / or modifier fragments to sentences in which the corresponding modifiers and / or modifier fragments can be inserted. For example, for common modifiers such as metaphors, adjectives, and adverbs, a metaphor recognition model and some manual rules can be used to construct training corpus pairs from specific types of unmodified sentences to sentences in which the corresponding modifier fragments can be inserted.
[0093] In addition, for paired training corpora of metaphor-modifying fragment types, the metaphor recognition model in the text correction strategy can be used to extract metaphor sentences from the massive text corpus, and metaphor words and boundary words can be extracted from the extracted metaphor sentences to form metaphor fragments. For example, metaphor words and boundary words such as "like..." and "as if..." in the metaphor sentences can be extracted. The sentences from which the metaphor fragments are extracted are treated as unmodified sentences. For paired training corpora of adjective and adverb types, it is only necessary to use natural language processing tools such as word segmentation, part-of-speech tagging, and syntactic analysis to extract the adjectives and adverbs that act as modifying components in the sentences, and then paired training corpora of adjective and adverb types can be constructed in the same way. It should be understood that in the embodiments of the present application, paired training corpora of various modification types are mixed together to train the second target network branch.
[0094] It should be understood that after obtaining the above-mentioned paired training corpora, a preset second network, such as a pointer generation network based on BART, can be trained based on the paired training corpora.
[0095] For example, Figure 8 As shown, Figure 8 This is a schematic diagram of the principle of the second target network branch. Figure 8 It can be seen that, assuming that the sentence included in the input text is: The scenery of Guilin is very beautiful. After the sentence is input into the second target network branch, the unmodified sentence obtained is X={x1, X2, ..., x n}, the modified fragment corresponding to the unmodified statement is Y = {y1, y2, ..., y m First, the unmodified sentence X is input into the BART encoder to obtain the vector representation of each character. Then, a pointer network is used to predict the position p where insertion can be performed. * , p * ∈[1, n]. Specifically, a pointer network is used to predict the position p where insertion can be performed. * The process can be expressed by the following formula:
[0096] H=BARTEncoder(X)
[0097] p * =Softmax(PointerNN(H))
[0098] Here, H represents the vector of the unmodified sentence X obtained after being encoded by the encoder BARTEncoder.
[0099] It should be understood that after predicting the possible insertion position p, the cross entropy loss between it and the actual insertion position p can be calculated to predict the insertion position. In the embodiment of the present application, the positioning loss loss can be used. p To represent the predicted insertion position. For example, loss p It can be expressed as:
[0100] loss p =CrossEntropy(p * , p)
[0101] It should be understood that for the insertable position p, the vector representation h after being encoded by the BART encoder BARTEncoder at the insertable position p is extracted. p As a guidance signal for subsequent generation, h p The vector representation H of each character on the encoder BARTEncoder side is sent to the BART encoder BARTEncoder to predict the first character y1 in the modified segment Y, and predict each character in the modified segment Y in turn. The corresponding calculation process can be expressed as follows:
[0102] h0=BARTDecoder(h p , H)
[0103] h t+1 =BARTDecoder(h t , h 0:t-1 , H)
[0104]
[0105]
[0106] In addition, it is necessary to use a multi-task joint loss to train the second target network branch (the pointer generation network based on BART). It should be understood that the multi-task joint loss function is the second loss function corresponding to the second target network branch, which is expressed as:
[0107] loss joint =(loss p +loss g ) / 2
[0108] Among them, h p represents the vector after being encoded by the BART encoder and can be inserted at position p; h0 represents h pThe vector obtained by the BART decoder; y t+1 , indicating the t+1th word in the modified segment; According to h t+1 The predicted word t+1; loss p Represents the loss of the task of predicting the insertion position; loss g Represents the loss of the task of generating modified fragments; loss joint represents the joint loss of the task of predicting the insertion position and generating the modified segment. Finally, after the second target network branch polishes the modifier, the output sentence is: The scenery of Guilin is as beautiful as a vivid painting.
[0109] S5023: Obtain a third preset number of third training samples, and train a preset third network architecture based on the third training samples to obtain the third target network branch.
[0110] Among them, the obtaining of the third preset number of third training samples includes: obtaining a third target number of target words, and constructing a synonym word graph of the target words; obtaining replacement words for each of the target words in the synonym word graph based on a preset search algorithm; replacing the target words in the target sentence with the corresponding replacement words in turn to obtain replacement sentences; obtaining the third preset number of replacement sentences, and using the replacement sentences as the third training samples.
[0111] In some embodiments, the training of the third network architecture based on the third training sample to obtain the third target network branch includes: inputting the third preset number of replacement sentences into the third network architecture, and calculating the perplexity of each replacement sentence respectively; obtaining the elegance of each replacement sentence based on a preset excellent text correction strategy; determining the end of training of the third network architecture according to the perplexity and the elegance, and obtaining the third target network branch.
[0112] In some embodiments, determining that the training of the third network architecture is completed based on the perplexity and the elegance to obtain the third target network branch includes: if the proportion of the perplexity less than the preset perplexity threshold is greater than a first preset proportion threshold, and the proportion of the elegance greater than the preset elegance threshold is greater than a second preset proportion threshold, then determining that the training of the third network architecture is completed and the third target network branch is obtained.
[0113] It should be noted that for a word w q And its corresponding interpretation P = {w1, w2, ..., w n}, calculate word w q Similarity to the word vector of each word in the interpretation, if the word wq If the similarity with some words in the definition is higher than a certain threshold, a line pointing to w q The directed edge connects the words in the explanation with w q For example, a synonym graph can be obtained by processing all the words in the Xinhua Dictionary.
[0114] After obtaining the synonym graph, some efficient search algorithms can be used to obtain some synonyms of the words to be upgraded in the synonym graph. Then, these synonyms are used as replacement words to replace the sentences where the words to be upgraded are located one by one to obtain several replaced sentences. Finally, the GPT model is used to calculate the perplexity of each sentence, and the excellent expression sentence recognition model is used to calculate the confidence of the excellent expression sentence of each sentence. Among them, the excellent expression sentence recognition model is a model in the text correction strategy. For example, the excellent expression sentence recognition model can be a binary classifier, which will calculate the probability of each sentence being an excellent expression sentence. Correspondingly, the calculated probability of an excellent expression sentence is the confidence of the excellent expression sentence.
[0115] It should be understood that only when there is only one replaced sentence whose perplexity and confidence of excellent expression are higher than the original sentence, the replacement word in this sentence will be used as an upgraded word to be upgraded. Figure 9 As shown, Figure 9 This is a schematic diagram of the principle of the third target network branch. Figure 9 The word to be upgraded is: persistence; according to the synonym word map of the word to be upgraded, the replacement words for the word to be upgraded include: perseverance, persistence, determination and maintenance.
[0116] In this embodiment, the training processes of the first target network branch, the second target network branch, and the third target network branch are independent of each other. It should be understood that in some other embodiments of the present application, the first target network branch, the second target network branch, and the third target network branch can also be jointly trained. Specifically, in the process of jointly training the first target network branch, the second target network branch, and the third target network branch, whether the training of each target network branch is completed can be evaluated based on the joint loss function. Among them, the joint loss function can be a weighted loss function of the first loss function, the second loss function, and the third loss function. In the weighted loss function, the first loss function, the second loss function, and the third loss function respectively correspond to preset weighting coefficients. Specifically, the weighted loss function can be represented by multiplying each loss function by its corresponding weighting coefficient and then adding them together. Among them, the weighting coefficients of each loss function can be set in advance and are not specifically limited here.
[0117] S503: Input the text to be processed into a pre-trained text polishing model, perform polishing on the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence, to obtain a target text.
[0118] From the above analysis, it can be seen that the text processing method provided in the embodiment of the present application first obtains a text to be processed, wherein the text to be processed includes at least one natural language sentence; then the text to be processed is input into a pre-trained text polishing model, and the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence are polished to obtain a target text. The text can be polished at the sentence, sentence modifier, and / or sentence vocabulary levels, which can efficiently and accurately improve the quality of the text content.
[0119] See also Figure 10 , Figure 10 1 is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0120] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause the processor to perform any computer method.
[0121] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0122] The internal memory provides an environment for the execution of computer programs in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any computer method.
[0123] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0124] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0125] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:
[0126] Acquire a text to be processed, where the text to be processed includes at least one natural language sentence;
[0127] The text to be processed is input into a pre-trained text polishing model, and the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence are polished to obtain a target text.
[0128] In one embodiment, before inputting the to-be-processed text into the pre-trained text polishing model, the method further includes: obtaining a preset number of training samples, training a preset language model, and obtaining the pre-trained text polishing model; wherein the text polishing model includes a first target network branch, a second target network branch, and a third target network branch that are independent of each other;
[0129] The first target network branch is used to detect natural language sentences that need to be polished in the text to be processed, and polish the natural language sentences that need to be polished to obtain a first target sentence;
[0130] The second target network branch is used to detect the target position that needs to be supplemented with a modifier in each of the natural language sentences, and supplement the modifier at the target position to obtain a second target sentence;
[0131] The third target network branch is used to detect the vocabulary to be optimized in each of the natural language sentences, and optimize the vocabulary to be optimized to obtain the target vocabulary.
[0132] In one embodiment, the preset language model includes: a first network architecture, a second network architecture, and a third network architecture; obtaining a preset number of training samples, training the preset language model, and obtaining the pre-trained text polishing model includes:
[0133] Obtaining a first preset number of first training samples, and training the first network architecture based on the first training samples to obtain the first target network branch;
[0134] Obtaining a second preset number of second training samples, and training the second network architecture based on the second training samples to obtain the second target network branch;
[0135] A third preset number of third training samples is obtained, and the third network architecture is trained based on the third training samples to obtain the third target network branch.
[0136] In one embodiment, obtaining a first preset number of first training samples includes:
[0137] Obtaining a first target amount of text data;
[0138] The text data is screened based on a preset excellent text correction strategy to obtain a first preset number of excellent sentences, and the excellent sentences are used as the first training samples.
[0139] In one embodiment, the training of a preset first network architecture based on the first training sample to obtain the first target network branch includes:
[0140] For any of the excellent sentences, extract keywords from the excellent sentence according to a preset keyword extraction algorithm to obtain at least one keyword;
[0141] Generating a first sequence of excellent sentences and a second sequence of keywords respectively;
[0142] Supplementing the second sequence to the first sequence to obtain a target sequence;
[0143] The target sequence corresponding to each excellent sentence is input into the first network architecture, and the first network architecture is trained based on the target sequence to obtain the first target network branch.
[0144] In one embodiment, training the first network architecture based on the target sequence to obtain the first target network branch includes:
[0145] Training the first network architecture based on the target sequence and detecting a value of a first loss function;
[0146] If the value of the first loss function is less than or equal to a first preset threshold, it is determined that the training of the first network architecture is completed, and the first target network branch is obtained.
[0147] In one embodiment, obtaining a second preset number of second training samples includes:
[0148] Acquire a second target amount of text data;
[0149] screening the text data based on a preset excellent text correction strategy to obtain a second preset number of excellent sentences;
[0150] Extracting modifiers from the excellent sentences to obtain sentences without modifiers;
[0151] A corpus pair is generated according to the sentence without the modifier and the first position of the modifier, and the corpus pair is used as the second training sample.
[0152] In one embodiment, the training of the preset second network architecture based on the second training sample to obtain the second target network branch includes:
[0153] For any pair of the corpus, analyzing the sentences without modifiers based on the second network architecture to obtain a word vector for each word in the sentences without modifiers;
[0154] Predicting a second position of the modifier in the word vector based on a preset pointer network;
[0155] Calculating the cross entropy between the first position and the second position, and determining the insertion position of the modifier according to the cross entropy;
[0156] Generate a position vector according to the insertion positions of all the modifiers;
[0157] Inputting the position vector and the word vector into a preset encoder to predict the modifier word at each insertion position;
[0158] The training of the second network architecture is determined to be completed according to the value of the second loss function, and the second target network branch is obtained.
[0159] In one embodiment, determining the end of training the second network architecture according to the value of the second loss function to obtain the second target network branch includes:
[0160] If the value of the second loss function is less than or equal to a preset threshold, it is determined that the training of the second network architecture is completed, and the second target network branch is obtained.
[0161] In one embodiment, obtaining a third preset number of third training samples includes:
[0162] Acquire a third target number of target words, and construct a synonym word graph for the target words;
[0163] Based on a preset search algorithm, obtaining a replacement word for each of the target words in the synonym word graph;
[0164] replacing the target words in the target sentence with the corresponding replacement words in sequence to obtain a replacement sentence;
[0165] A third preset number of replacement sentences are obtained, and the replacement sentences are used as the third training samples.
[0166] In one embodiment, the training of a preset third network architecture based on the third training sample to obtain the third target network branch includes:
[0167] Inputting the third preset number of replacement sentences into the third network architecture, and calculating the perplexity of each replacement sentence respectively;
[0168] Obtaining the beauty of each replacement sentence based on a preset excellent text correction strategy;
[0169] The training of the third network architecture is determined to be completed according to the perplexity and the elegance, and the third target network branch is obtained.
[0170] In one embodiment, determining the end of training the third network architecture according to the perplexity and the elegance to obtain the third target network branch includes:
[0171] If the proportion of the perplexity less than the preset perplexity threshold is greater than a first preset proportion threshold, and the proportion of the elegance greater than the preset elegance threshold is greater than a second preset proportion threshold, it is determined that the training of the third network architecture is completed and the third target network branch is obtained.
[0172] In one embodiment, the first network architecture is a keyword-guided GPT sentence generation network model;
[0173] The second network architecture is a pointer generation network model based on BART; the third network architecture is a GPT vocabulary generation network model based on replacement sentence guidance.
[0174] In an embodiment of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement the present application. Figure 1The illustrated embodiments provide a computer method.
[0175] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.
[0176] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A text processing method, characterized in that: The method comprises: Acquire a text to be processed, where the text to be processed includes at least one natural language sentence; Obtain a preset number of training samples, train a preset language model, and obtain a pre-trained text polishing model; Inputting the text to be processed into a pre-trained text polishing model, performing polishing processing on the natural language sentence, the modifiers in the natural language sentence, and / or the vocabulary to be optimized in the natural language sentence, to obtain a target text; The pre-trained text polishing model includes a first target network branch and a second target network branch, the preset language model includes a first network architecture and a second network architecture, and the pre-trained text polishing model is obtained by obtaining a preset number of training samples and training the preset language model, including: For any excellent sentence in a first preset number of first training samples, extract keywords from the excellent sentence according to a preset keyword extraction algorithm to obtain at least one keyword; generate a first sequence of the excellent sentence and a second sequence of the keywords respectively; append the second sequence to the first sequence to obtain a target sequence; input the target sequence corresponding to each excellent sentence into the preset language model, and train the first network architecture based on the target sequence to obtain the first target network branch; and Analyze any pair of corpora in a second preset number of second training samples to obtain a word vector for each word in the sentence without modifiers in the corpus pair, and predict the second position of the modifier in the word vector based on the first position of the modifier in the corpus pair; predict the insertion position of the modifier based on the first position and the second position; generate a position vector according to the insertion positions of all the modifiers, and determine the predicted modifier for each insertion position based on the word vector of each word and the position vector; train the second network architecture based on the modifiers and the predicted modifiers to obtain the second target network branch.
2. The method according to claim 1, characterized in that The text polishing model includes the first target network branch, the second target network branch and the third target network branch which are independent of each other; The first target network branch is used to detect natural language sentences that need to be polished in the text to be processed, and polish the natural language sentences that need to be polished to obtain a first target sentence; The second target network branch is used to detect the target position that needs to be supplemented with a modifier in each of the natural language sentences, and supplement the modifier at the target position to obtain a second target sentence; The third target network branch is used to detect the vocabulary to be optimized in each of the natural language sentences, and optimize the vocabulary to be optimized to obtain the target vocabulary.
3. The method according to claim 2, characterized in that The preset language model further includes: a third network architecture; The step of obtaining a preset number of training samples and training a preset language model to obtain the pre-trained text polishing model further includes: A third preset number of third training samples is obtained, and the third network architecture is trained based on the third training samples to obtain the third target network branch.
4. The method according to claim 1, wherein The obtaining of a first preset number of first training samples includes: Acquire a first target amount of text data; The text data is screened based on a preset excellent text correction strategy to obtain a first preset number of excellent sentences, and the excellent sentences are used as the first training samples.
5. The method according to claim 1, characterized in that The training of the first network architecture based on the target sequence to obtain the first target network branch includes: Training the first network architecture based on the target sequence and detecting a value of a first loss function; If the value of the first loss function is less than or equal to a first preset threshold, it is determined that the training of the first network architecture is completed, and the first target network branch is obtained.
6. The method according to claim 1, characterized in that The process of acquiring the second preset number of second training samples includes: Acquire a second target amount of text data; screening the text data based on a preset excellent text correction strategy to obtain a second preset number of excellent sentences; Extracting modifiers from the excellent sentences to obtain sentences without modifiers; A corpus pair is generated according to the sentence without the modifier and the first position of the modifier, and the corpus pair is used as the second training sample.
7. The method according to claim 6, characterized in that The training of the second network architecture based on the second training sample to obtain the second target network branch includes: For any pair of the corpus, analyzing the sentences without modifiers based on the second network architecture to obtain a word vector for each word in the sentences without modifiers; Predicting a second position of the modifier in the word vector based on a preset pointer network; Calculating the cross entropy between the first position and the second position, and determining the insertion position of the modifier according to the cross entropy; Generate a position vector according to the insertion positions of all the modifiers; Inputting the position vector and the word vector into a preset encoder to predict the modifier word at each insertion position; The training of the second network architecture is determined to be completed according to the value of the second loss function, and the second target network branch is obtained.
8. The method according to claim 7, characterized in that The step of determining the end of training the second network architecture according to the value of the second loss function to obtain the second target network branch includes: If the value of the second loss function is less than or equal to a preset threshold, it is determined that the training of the second network architecture is completed, and the second target network branch is obtained.
9. The method according to claim 3, characterized in that The obtaining of a third preset number of third training samples includes: Acquire a third target number of target words, and construct a synonym word graph for the target words; Based on a preset search algorithm, obtaining a replacement word for each of the target words in the synonym word graph; replacing the target words in the target sentence with the corresponding replacement words in sequence to obtain a replacement sentence; A third preset number of replacement sentences are obtained, and the replacement sentences are used as the third training samples.
10. The method according to claim 9, characterized in that The training of the third network architecture based on the third training sample to obtain the third target network branch includes: Inputting the third preset number of replacement sentences into the third network architecture, and calculating the perplexity of each replacement sentence respectively; Obtaining the beauty of each replacement sentence based on a preset excellent text correction strategy; The training of the third network architecture is determined to be completed according to the perplexity and the elegance, and the third target network branch is obtained.
11. The method according to claim 10, characterized in that The step of determining, based on the perplexity and the elegance, that the training of the third network architecture is completed and obtaining the third target network branch comprises: If the proportion of the perplexity less than the preset perplexity threshold is greater than a first preset proportion threshold, and the proportion of the elegance greater than the preset elegance threshold is greater than a second preset proportion threshold, it is determined that the training of the third network architecture is completed and the third target network branch is obtained.
12. The method according to claim 3, characterized in that The first network architecture is a GPT sentence generation network model based on keyword guidance; The second network architecture is a pointer generation network model based on BART; The third network architecture is a GPT vocabulary generation network model guided by replacement sentences.
13. A computer device, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the steps of the text processing method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the steps of the text processing method according to any one of claims 1 to 12.
Citation Information
Patent Citations
A writing aiding method based on neural network
CN109522411A
Text creation method, device, equipment, and storage medium
CN111832275A
Text processing method and device, electronic equipment and medium
CN112115706A
Method and device for realizing natural language understanding in human-computer interaction system
CN112632962A