Abstraction type text summarization method and system based on semi-supervised learning
By adopting the semi-supervised learning TST-CycleGAN model in text digest generation, combining Seq2Seq structure and Bi-LSTM, adding attention mechanism and TextRank algorithm, the problems of semantic incoherence and insufficient generalization capabilities in text digest generation are solved, and high-quality and highly readable text digest generation are achieved.
Patent Information
- Application Number
- CN202510217111.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively capture long-distance text dependencies in text summary generation, resulting in incoherent semantics of the generated summary. In a low-resource environment, the model generalization ability is insufficient, and overfitting is prone to occur.
The TST-CycleGAN model based on semi-supervised learning is adopted, combined with the Seq2Seq structure and the bidirectional long and short-term memory network (Bi-LSTM), and the attention mechanism and TextRank algorithm are added to enhance semantic information extraction. Through the optimization generator of cyclic consistency loss and identity mapping loss, the generated digest is ensured to be complete and readable.
It improves the semantic coherence and readability of text summary, enhances the generalization ability of the model, reduces the cost of manual labeling, alleviates the overfitting problem, and makes the generated abstract more stable in texts in different fields and styles.
Smart Images

Figure CN120144754A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an abstract text summarization method based on semi-supervised learning, belonging to the technical field of abstract text summarization. Background Art
[0002] With the continuous development of the information society, people are exposed to a large amount of text content through various apps every day, and the long and complex information is unacceptable to people for a while. Therefore, it is necessary to establish a system that can provide a concise description of the full text while still conveying the key information in the original text. This task is usually called automatic text summarization.
[0003] Text summarization enables users to easily obtain the key information and overall meaning of a document without reading the entire content. Text summarization in the field of natural language processing (NLP) is divided into extractive summarization and abstractive summarization. Extractive summarization generates a summary by extracting important phrases or sentences from the original text, and usually these extracted fragments directly constitute the final summary. Extractive summarization models tend to learn important phrases or text parts in the training data, which are used to represent significant pattern features and generate fluent summaries. In contrast, abstractive summarization models aim to generate new, shorter text content that fully reflects the key and significant information of the original text. Abstractive summarization not only needs to generate a natural and linguistically coherent summary, but also must be combined with complex natural language processing mechanisms to accurately understand and compress the information in the original document. Therefore, abstractive text summarization is generally considered more challenging than extractive summarization.
[0004] With the development of deep learning technology, many methods based on Seq2Seq, RNN (Recurrent Neural Network), and LSTM (Long Short-Term Memory) models have been widely used. The Seq2Seq (Sequence-to-Sequence) model is a deep learning model based on the encoder-decoder structure and is widely applied to sequence-to-sequence tasks. Seq2Seq was initially designed for machine translation tasks, but later researchers extended it to the field of text summarization. In this model, the Encoder is responsible for converting the input text sequence into a special sequence that the model can process, extracting the semantic information of the text, and then passing this information to the Decoder. The Decoder generates the corresponding summary based on the special sequence received from the encoder and converts it into a text sequence for output. In the Encoder-Decoder structure, variants such as RNN and LSTM are widely used in this process. When the RNN model processes text data, it treats it as a continuous time series, extracts the interdependencies between sequences to capture hidden semantic information. However, traditional RNNs faced the problems of vanishing gradients and exploding gradients in early experiments, which would cause the model to be unable to capture the dependencies of long-distance text sequences. When generating text summaries, it may lose information over long distances, resulting in incoherent context semantics. In a unidirectional LSTM, although the problems of vanishing gradients and explosions are effectively alleviated through the gating mechanism and activation functions, the model can still only utilize the information from the "previous context" because the direction of the input text is single. In practical applications, in order to make accurate predictions, it is often necessary to synthesize the information of the entire sequence. Therefore, the current mainstream method is to use the Bidirectional Long Short-Term Memory network (Bi-LSTM). Bi-LSTM adds a layer of neurons for receiving the reverse sequence based on the traditional LSTM, thus combining the information from both directions. In this way, both the semantic information of the previous context and that of the subsequent context are obtained, and the final output result of the Bi-LSTM layer is obtained after a series of processing of the semantic information before and after.
[0005] Generative Adversarial Networks (GANs) were initially proposed in the field of computer vision. However, with the development of deep learning technologies, GANs have gradually demonstrated their potential in the text domain. The core idea of GANs is to improve the quality of generated text through a game between a generator and a discriminator. In the text domain, the generator is responsible for generating new text content, while the discriminator judges the similarity between the generated text and the real text. This process is iterated until the generator can deceive the discriminator into being unable to distinguish between the generated text and the real text. Although GANs have achieved great success in the image domain, they face some unique challenges in the text domain. For example, text is discrete data, while GANs usually deal with continuous data, which makes it impossible for the generator to be optimized through conventional gradient backpropagation in text generation. To address the many problems of GAN models in the text domain, researchers have developed a series of GAN models suitable for the text domain, such as SeqGAN and TextGAN, as Figure 1 shown. Therefore, generative adversarial networks have been applied to various tasks in the text domain. It can generate or verify true and false data in a differentiable manner through multi-task training objectives to produce the expected output. In the text summarization task, the purpose of the generator is to generate a summary for the input article, and the training objective is to generate a summary that looks like it was written by a human to deceive the discriminator. Nowadays, in the text summarization task, excellent GAN models adopt a Seq2Seq structure generator based on bidirectional LSTM (Bi-LSTM) as the encoder and decoder. This generator can better capture the context information of the input text and generate a coherent and accurate summary.
[0006] CycleGAN is a generative adversarial network (GAN) architecture for image-to-image translation tasks. This method is widely used in multiple fields, including but not limited to image style transfer, domain adaptation, and super-resolution. The core idea of CycleGAN is to train between two pairs of generator-discriminator, with each pair consisting of a generator and a discriminator. One generator is responsible for converting images from the source domain to the target domain, and the other generator is responsible for converting the images back from the target domain to the source domain. The task of the discriminator is to distinguish between the generated images and the real images in the corresponding domain. By introducing a cycle consistency loss, it is ensured that the generated images can be converted back to their original form. Figure 2 illustrates the principle of CycleGAN in one direction. The highlight of CycleGAN is that in addition to the adversarial loss, there are also identity mapping loss and cycle consistency loss. The identity mapping loss means that it aims to ensure that when the generator processes data that already belongs to the target domain, it should not make unnecessary changes to the data. The cycle consistency loss ensures that the converted data can be restored to the original data.
[0007] CycleGAN has been widely applied in the field of images and text style transfer tasks. In the field of text summarization, the article and the summary can be regarded as the source domain or the target domain respectively to complete the generation of text summaries. Prior to this, some researchers have been inspired by CycleGAN and referred to its cyclic consistency idea for research, but they did not conduct experiments using the GAN architecture. Summary of the Invention
[0008] Aiming at the deficiencies of the prior art, the present invention provides an abstractive text summarization method based on semi-supervised learning;
[0009] The present invention designs a TST-CycleGAN model using a complete GAN architecture and applies it to the text summarization task. The core idea of CycleGAN is to establish two pairs of generator-discriminator respectively for converting the article into a summary and converting the summary into an article, as shown in Figures 3(a) and 3(b), where, G AS is the generator that generates the summary from the article, and F SA is the generator that reversely generates the article from the summary.
[0010] The generator of the present invention adopts a Seq2Seq structure, in which the encoder adopts a bidirectional long short-term memory network (Bi-LSTM), and the decoder adopts an LSTM with an attention mechanism added. The encoder adopting Bi-LSTM can capture the context dependencies of the input sequence, so as to better understand the overall input. Especially in text processing, the relationship of the context has an important impact on the semantics of each word. The main task of the decoder is to generate the target summary according to the input sequence. Since the generation process is carried out step by step, the generation at the current time step only depends on the past context and does not require future information. Therefore, a unidirectional LSTM is more suitable. With the addition of the attention mechanism, the decoder can flexibly focus on the output of the encoder when generating each step, ensuring that each part generated is based on appropriate context.
[0011] Considering the importance of the semantic information of the article for the text summarization task, the present invention also adds a module for enhancing semantic information to this model. This module uses the TextRank algorithm to extract keywords (keywords are words that appear frequently in the article and are closely related to other important words, and they are usually the core content of the text or the words that summarize the central idea of the text), and then concatenates the extracted keywords with the text of the original article as the input of the encoder to help the model better understand the core content of the article during the generation process. By processing the article and keywords simultaneously, the encoder can provide richer context information for the decoder.
[0012] The present invention provides an abstractive text summarization system based on semi-supervised learning;
[0013] Glossary:
[0014] 1. The TextRank algorithm is a graph-based ranking algorithm mainly used for text processing tasks such as keyword extraction and text summarization. It draws on the idea of PageRank and evaluates the importance of text units (such as words or sentences) by constructing a graph model. Its core idea is divided into three points:
[0015] ① Graph construction: Text units (such as words or sentences) are used as nodes in the graph, and the edges between nodes represent their relationships (such as co-occurrence or similarity).
[0016] ② Weight calculation: The weights of each node are calculated iteratively, and nodes with high weights are considered more important.
[0017] ③ Iterative update: The weights of nodes are dynamically updated according to the weights of their neighbor nodes until convergence.
[0018] The detailed steps of the TextRank algorithm are divided into the following five points:
[0019] ① Text preprocessing: including word segmentation, stop word removal, etc.
[0020] ② Graph construction: Construct a graph based on the relationships of text units.
[0021] ③ Weight initialization: Assign initial weights to each node.
[0022] ④ Iterative calculation: Update the node weights iteratively until stable.
[0023] ⑤ Result output: Sort according to the weights and extract keywords or generate a summary.
[0024] 2. The LCSTS dataset. LCSTS is a large-scale Chinese short text summarization dataset that has collected more than 2 million Chinese short texts published on the Weibo website and their corresponding summaries. The data are all generated by real users, and these summaries are concise summaries given by the text authors. Each sample in the dataset consists of a corresponding short text and a summary, which is suitable for the application of text summarization tasks. In addition, 10,000 pairs of short summaries and their texts in the dataset are manually marked for evaluating the correlation between the two.
[0025] The technical solution of the present invention is as follows:
[0026] An abstractive text summarization method based on semi-supervised learning, including:
[0027] Obtain data and perform preprocessing;
[0028] Input the preprocessed data into the trained TST-CycleGAN model to achieve automatic text summarization.
[0029] Preferably according to the present invention, in the training of the TST-CycleGAN model, the training data set is the LCSTS data set.
[0030] Preferably according to the present invention, the TST-CycleGAN model includes two pairs of generator-discriminators; one pair of generator-discriminators is used to convert an article into an abstract, and the other pair of generator-discriminators is used to convert an abstract into an article;
[0031] In the generator-discriminator, the generator adopts a Seq2Seq structure. The generator includes an encoder and a decoder. The encoder adopts a bidirectional long short-term memory network, and the decoder adopts an LSTM with an attention mechanism; the encoder is used to capture the context dependence of the input sequence, and the decoder is used to generate a target abstract according to the input sequence;
[0032] The discriminator includes a three-layer convolutional neural network, which is used to extract local features in the abstract;
[0033] Preferably according to the present invention, the TST-CycleGAN model includes: a generator G for text to generate an abstract AS , a generator F for abstract to generate text SA , a discriminator D for distinguishing true abstracts and false abstracts 1 and a discriminator D for distinguishing true texts and false texts 2 ; in the training of the TST-CycleGAN model,
[0034] First, perform a warm-up session of supervised learning; optimize the supervised learning through reinforcement learning, input paired text and abstract sequences so that the CycleGAN model learns the mapping relationship between the text and abstract sequences, and initially master the rules for generating abstracts;
[0035] Second, perform an unsupervised learning training session; including:
[0036] First, input the original text a into the module for enhancing semantic information, extract keywords using the TextRank algorithm and splice them with the original text to obtain text a ′ , and then text a ′ and the original abstract s sequence are respectively input into two generators G AS and F SA , to generate machine abstracts and machine texts;
[0037] Subsequently, the original text a and the original abstract s sequence, as well as the generated machine abstracts and machine texts, are respectively input into discriminator D a and discriminator D S , and calculate the adversarial loss;
[0038] Then, the generated machine abstract and machine text are respectively input into F SA and G AS , to generate fake text a fake and fake abstract s fake ;
[0039] Finally, the cycle consistency loss is used to calculate the errors between a and a fake , s and s fake respectively. Meanwhile, the identity mapping loss is used to constrain the generator, such that when the original text a is directly input into the generator F SA , the output is a, and when the original abstract s is directly input into the generator G AS , the output is s.
[0040] Preferably according to the present invention, the TST-CycleGAN model further includes a module incorporating the enhanced semantic information, i.e., the TextRank algorithm;
[0041] First, after the article is input into the module with enhanced semantic information, text preprocessing operations are performed, including word segmentation and stop word removal; the BERT model is used to extract the semantic information of the text sequence, and a lexical graph is constructed by calculating semantic similarity, with each word regarded as a node in the graph;
[0042] Then, the edges of the graph are established based on the co-occurrence relationship of the words; the co-occurrence relationship is based on the window size, and when two words appear in the window simultaneously, an edge is established between these two words;
[0043] Next, the initial score of each node is set to 1, and then the weight of each node is iteratively calculated through Equation (1):
[0044]
[0045] In Equation (1), S(V i ) represents the score of node V i , d is the damping factor, In(V i ) and Out(V + ) respectively represent the input nodes and output nodes connected to node V i , V i , V + respectively represent two different nodes;
[0046] After iterative calculation, the scores of each node are sorted from large to small, and the top n words with the highest scores are selected as keywords;
[0047] Finally, these keywords are encoded and concatenated with the text sequence to generate a new text sequence;
[0048] Further preferably, the semantic information of the text sequence is extracted by using the BERT model, and a lexical graph is constructed by calculating semantic similarity, including:
[0049] The text sequence x = (x 1 , x 2 ,..., x / ) is input into the BERT model to obtain the context semantic representation of each word in the text sequence x, as shown in Equation (2):
[0050] h i = BERT(x 5 )(2);
[0051] In Equation (2), x 5 represents the i-th word in the text, and h i represents the high-dimensional word vector of x 5 ;
[0052] The semantic similarity between two nodes is calculated using cosine similarity, as shown in Equation (3):
[0053]
[0054] In Equation (3), and respectively represent two vectors. The numerator represents the dot product of and , and the denominator represents the modulus normalization of and . The value range of the semantic similarity is between [-1, 1].
[0055] Preferably according to the present invention, the loss function of the TST-CycleGAN model adopts the cross-entropy loss function, as shown in Equation (4):
[0056]
[0057] In Equation (4), L(a) represents the cross-entropy loss, a represents a single text, and p i (a) represents the probability of the true class i, represents the probability of the model predicting class i.
[0058] Preferably according to the present invention, the training process of the TST-CycleGAN model is as follows:
[0059] After the article is input into the TST-CycleGAN model, the identity mapping loss L SA (a) is first obtained through the generator F iM , and then passes through the generators G AS and F SAGenerate fake articles in a loop to obtain the cycle consistency loss L NONPe (a);
[0060] Similarly, after inputting the abstract into the TST-CycleGAN model, L is obtained iM (s) and L NONPe (s), and finally, the weighted sum is performed with the adversarial loss L of the GAN architecture aMD to obtain respectively:
[0061] Among them, L iM (s) and L iM (a) are respectively the identity mapping losses of the abstract and the text; L NONPe (s) and L NONPe (a) are respectively the cycle consistency losses of the abstract and the text; L aMD (F SA ) and L aMD (G AS ) are respectively the adversarial losses of the two generators.
[0063] According to the preference of the present invention, the warm-up link of supervised learning specifically includes:
[0064] The first step is to input the text sequence of the real text and the real abstract label into both the generator G AS and F SA , and calculate the losses L AS and L SA output by the generators G Q and F Y through the forward propagation calculation of the BiLSTM encoding-decoding structure (Encoder-Decoder);
[0065] The second step is to optimize the generator using the policy gradient method in reinforcement learning: input the real abstract into the discriminator D 1 to obtain the prediction result, and use the prediction result as the reward signal, then calculate the logarithmic probability of this reward signal to calculate the policy gradient loss; finally, multiply the mean of the reward signal and the mean of the logarithmic probability to obtain the policy gradient loss;
[0066] Optimize the generators G AS and F SA respectively as shown in Equations (5) and (6):
[0067]
[0068] In Equations (5) and (6), θ is the weight parameter; are respectively the policy gradient losses for calculating the abstract and the text; LQ and L Y is the loss of the forward propagation output of the BiLSTM generator.
[0069] Preferably according to the present invention, the training process of unsupervised learning specifically includes:
[0070] First, the original text is input into the module for enhancing semantic information to obtain the enhanced text, and then the enhanced text and the original summary are respectively input into the tokenizer to obtain the text sequence a = (a 1 , a 2 ,..., a / ) and the summary sequence s = (s 1 , s 2 ,..., s / );
[0071] Then, a and s are respectively input into the generators G AS and F SA , to generate the machine summary sequence s ′ = (s 1 ′ , s ′ 2 ,..., s ′ / ) and the machine text sequence a ′ = (a ′ 1 , a ′ 2 ,..., a ′ / );
[0072] Finally, the generated machine summary sequence s ′ and the machine text sequence a ′ are respectively input into F SA and G AS , to generate fake texts a fake and fake summaries s fake similar to the original text and summary;
[0073] Perform processing of the loss function, and use the cross - entropy loss function to calculate the cycle - consistency loss L fake Ua, a NONPe_a V of the fake text a fake V and the cycle - consistency loss L fake NONPe_` (s, s(s, s fake ) of the fake summary s
[0074] Preferably according to the present invention, in the training process of unsupervised learning, an identity mapping loss is introduced, including:
[0075] First, input the text sequence a = (a 1 , a 2 , …, a / ) and the abstract sequence s = (s 1 , s 2 , …, s / ) into the generators F SA and G AS respectively to obtain the output sequences a iM and s iM ;
[0076] Subsequently, calculate the identity mapping loss and respectively to optimize the generators F SA and G AS .
[0077] The loss function for optimizing the generators includes cyclic consistency loss, identity mapping loss, and adversarial loss, and the formula is as follows:
[0078]
[0079] where α and β are weight coefficients; L NONPe = ωL NONPe_a + ωL NONPe_` , ω is a weight coefficient; L NONPe , respectively refer to the cyclic consistency loss, the identity mapping loss of the abstract, and the adversarial loss for the reverse optimization of the discriminator D 1 , is the identity mapping loss of the text, is the adversarial loss for the reverse optimization of the discriminator D 2 , which is calculated through the cross - entropy loss function; refers to the loss function of the generator G AS , refers to the loss function of the generator F SA .
[0080] Preferably according to the present invention, during the training process of the discriminator,
[0081] first, use the generators to generate fake abstracts and fake texts, and use the Jieba tokenizer to tokenize the real text and real abstract;
[0082] then use the discriminators D 1 and D 2 composed of a three - layer convolutional neural network to generate prediction results, and this prediction result is the predicted value after softmax processing;
[0083] Next, set the real text and abstract labels to 1, and the fake text and abstract labels to 0, and calculate the loss L between the prediction result of the previous step and the true / false labels. g ; The formula is as follows:
[0084]
[0085] where γ is the weight coefficient; and are the losses between the predicted values of true / false texts and the true / false labels respectively.
[0086] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of an abstractive text summarization method based on semi-supervised learning.
[0087] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of an abstractive text summarization method based on semi-supervised learning.
[0088] An abstractive text summarization system based on semi-supervised learning includes:
[0089] A data acquisition and preprocessing module, configured to: acquire data and perform preprocessing;
[0090] An automatic text summarization module, configured to: input the preprocessed data into the trained TST-CycleGAN model to implement automatic text summarization.
[0091] The beneficial effects of the present invention are:
[0092] 1. Semi-supervised learning can utilize a large amount of unlabeled text data to help the model learn more comprehensive text features and semantic information, and thus generate more coherent, readable, and information-covered summaries.
[0093] 2. Labeling high-quality text summaries requires a large amount of manual input, but semi-supervised learning can reduce the manual labeling cost.
[0094] 3. Semi-supervised learning can enhance the generalization ability of the summary model. In a low-resource environment, semi-supervised learning can alleviate the overfitting problem and make the model perform more stably on texts in different domains and styles. Description of the Drawings
[0095] Figure 1 is a simple structural schematic diagram of the GAN model;
[0096] Figure 2 is a schematic diagram of the working principle of the CycleGAN model;
[0097] Figure 3(a) is a schematic diagram of generating an abstract for an article;
[0098] Figure 3(b) is a schematic diagram of generating an article from an abstract;
[0099] Figure 4(a) is a schematic diagram of identity mapping loss;
[0100] Figure 4(b) is a schematic diagram of cycle consistency loss;
[0101] Figure 5 It is a schematic diagram of improving the TextRank algorithm for the BERT model. Specific implementation manners
[0102] The present invention will be further limited below in conjunction with the specification drawings and embodiments, but not limited thereto.
[0103] Embodiment 1
[0104] An abstractive text summarization method based on semi-supervised learning includes:
[0105] Obtain data and perform preprocessing;
[0106] Input the preprocessed data into the trained TST-CycleGAN model to achieve automatic text summarization.
[0107] In the present invention, data acquisition and preprocessing are key links in the automatic text summarization task. Data usually refers to text corpora for training and testing, including news articles, academic papers, travel guides, user reviews, etc., which can come from various sources such as public data sets (such as LCSTS, CSL), enterprise internal documents, API crawled data, etc. In the data preprocessing process, first, HTML tags need to be removed, special characters need to be removed, and sentence segmentation is performed. Natural language processing (NLP) techniques are used to perform operations such as word segmentation, stop word removal, and part-of-speech tagging on the text. In addition, key sentences are extracted through TextRank or TF-IDF to enhance the summary quality. Finally, all preprocessed texts are converted into a standardized data format (such as JSON, CSV) for easy input into the training model for summary generation.
[0108] The TST-CycleGAN automatic text summarization method proposed by the present invention is a semi-supervised learning model based on the CycleGAN structure. First, in the model training stage, a small part of the data with manually annotated summaries is used for supervised learning to enable the model to master the correspondence between the text and the summary. Then, in the unsupervised learning stage, the model uses a large amount of unannotated text and is optimized through cycle consistency loss, identity mapping loss, and adversarial loss, enabling it to generate high-quality summaries even without manually annotated summaries. In practical applications, the preprocessed text is input into the trained TST-CycleGAN model, and through the adversarial learning process of the bidirectional generator (text → summary, summary → original text), a concise and key information-containing summary is finally generated, ensuring its semantic integrity and readability.
[0109] Embodiment 2
[0110] A semi-supervised learning-based abstractive text summarization method according to Embodiment 1, wherein:
[0111] In the training of the TST-CycleGAN model, the training data set is the LCSTS data set.
[0112] The TST-CycleGAN model includes two pairs of generator-discriminators; one pair of generator-discriminators is used to convert the article into a summary, and the other pair of generator-discriminators is used to convert the summary into an article;
[0113] In the generator-discriminator, the generator adopts a Seq2Seq structure. The generator includes an encoder and a decoder. The encoder adopts a bidirectional long short-term memory network (Bi-LSTM), and the decoder adopts an LSTM with an attention mechanism; the encoder is used to capture the context dependence of the input sequence, so as to better understand the overall input. Especially in text processing, the relationship of the context has an important impact on the semantics of each word. The decoder is used to generate the target summary according to the input sequence; the attention mechanism helps the decoder to more accurately focus on the key information in the text when generating the summary, thereby improving the quality and coherence of the summary. Since the generation process is carried out step by step, the generation at the current time step only depends on the past context and does not require future information. Therefore, a unidirectional LSTM is more suitable.
[0114] The discriminator includes a three-layer convolutional neural network (CNN), which is used to effectively extract local features in the summary; helping the discriminator to more accurately distinguish between the generated summary and the real summary.
[0115] The framework of the present invention innovatively applies the GAN model in the field of text summarization to improve the CycleGAN model, proposes the TS-CycleGAN (TextSummarization-CycleGAN) model, and uses semi-supervised learning for model training. In the TS-CycleGAN model, the unique cycle consistency loss and identity mapping loss are combined with the adversarial loss unique to the GAN series. Two generators are trained separately through backpropagation, while the discriminator is trained in the same way as the traditional GAN model.
[0116] The TST-CycleGAN model includes: a generator G for generating summaries from text AS , a generator F for generating text from summaries SA , a discriminator D for distinguishing true summaries from false summaries 1 and a discriminator D for distinguishing true text from false text 2 ; In the training of the TST-CycleGAN model,
[0117] Due to the unsupervised learning characteristics of the CycleGAN model and the particularity that text summarization technology cannot be fully unsupervised, in the first step, a warm-up session of supervised learning is carried out; the classical supervised learning is optimized through reinforcement learning, and paired text and summary sequences are input to enable the CycleGAN model to learn the mapping relationship between the text and summary sequences and initially master the rules for generating summaries;
[0118] In the second step, an unsupervised learning training session is carried out; including:
[0119] First, the original text a is input into the module for enhancing semantic information, and keywords are extracted using the TextRank algorithm and concatenated with the original text to obtain text a ′ , and then text a ′ and the original summary s sequence are respectively input into the two generators G AS and F SA to generate machine summaries and machine texts;
[0120] Subsequently, the original text a and the original summary s sequence, as well as the generated machine summaries and machine texts, are respectively input into the discriminator D A and the discriminator D S , and the adversarial loss is calculated;
[0121] Then, the generated machine summaries and machine texts are respectively input into F SA and G AS to generate fake texts a fake and fake summaries s fake that restore the original text and the original summary;
[0122] Finally, the cyclic consistency loss is used to calculate the errors between a and a fake and between s and s fake respectively. Meanwhile, the identity mapping loss is used to constrain the generator, such that when the original text a is directly input into the generator F SA , the output is a, and when the original summary s is directly input into the generator G AS , the output is s. This prevents the model from overly modifying the input text and ensures the stability and semantic integrity of the generated summary.
[0123] All three losses adopt the cross-entropy loss function.
[0124] The present invention employs two pairs of generator-discriminator: (G AS , D S ) for generating summaries, and (F SA , D A ) for generating articles. The purpose of this design is to combine the cyclic consistency loss and the identity mapping loss to enhance the model's generation ability. In the text summarization task, the cyclic consistency loss ensures that when the summary generator converts the original text into a concise summary, key semantic information is not lost; meanwhile, when the generated summary is reversely converted into a long text, it should be able to maintain semantic consistency with the original long text. This mechanism guarantees that the generated summary is not only concise but also can accurately convey the core idea of the original text. The identity mapping loss ensures that when the input text is already highly concise or meets the target summary standard, the generator will not further simplify or modify it. This mechanism helps the model avoid redundant processing of the optimized input and ensures that the generator only performs conversion operations when necessary. Figures 4(a) and 4(b) respectively show the applications of the cyclic consistency loss and the identity mapping loss in this framework.
[0125] First, after adding a module for enhancing semantic information to the article input, text preprocessing operations are performed, including word segmentation and stop word removal; since the TextRank algorithm relies on the relationships between words when extracting keywords, before constructing the graph structure, it is necessary to use the BERT model to extract the semantic information of the text sequence, and a lexical graph is constructed by calculating semantic similarity, with each word regarded as a node in the graph;
[0126] Then, the edges of the graph are established based on the co-occurrence relationship of the words; the co-occurrence relationship is based on the window size (usually set to 5, that is, the current word is related to the 5 words before and after it), and when two words appear in the window simultaneously, an edge is established between these two words; the weight of the edge is calculated according to the co-occurrence frequency, or different weights are assigned by other means (such as semantic similarity), or even the same weight is assigned to each edge;
[0127] Next, the initial score of each node is set to 1, and then the weight of each node is iteratively calculated through Equation (1):
[0128]
[0129] In formula (1), S(V i ) represents the node V i score, d is the damping factor (usually set to 0.85), In(V i ) and Out(V + ) respectively represent the nodes V i The connected input and output nodes, V i 、V + Represent two different nodes respectively;
[0130] After iterative calculation, sort the scores of each node from large to small, and select the top n words as keywords;
[0131] Finally, these keywords are encoded and concatenated with the text sequence to generate a new text sequence; the text is input into the word segmenter to obtain the word x. 5 The text sequence x=(x 1 ,x 2 ,...,x / ), word segmenters include jieba Chinese word segmenter, word segmenter built into BERT model, etc.; about concatenation, for example, keyword: artificial intelligence. Artificial intelligence (AI), as one of the most cutting-edge technologies today, is gradually penetrating into people's daily lives, such as smart assistants, autonomous driving, medical diagnosis and many other fields.
[0132] The BERT model is used to extract semantic information from text sequences and to construct a vocabulary graph by calculating semantic similarity. This includes:
[0133] The text sequence x=(x 1 ,x 2 ,...,x / ) is input into the BERT model to obtain the contextual semantic representation of each word in the text sequence x, as shown in formula (2):
[0134] h i =BERT(x 5 )(2);
[0135] In formula (2), x 5 represents the i-th word in the text, h i Represents x 5 High-dimensional word vectors;
[0136] The semantic similarity between two nodes is calculated using cosine similarity, as shown in formula (3):
[0137]
[0138] In formula (3), and respectively represent two vectors, and the numerator represents and 's dot product, and the denominator represents and 's modulus normalization. The value range of semantic similarity is between [-1, 1] (-1 indicates the opposite, and 1 indicates complete correlation). Since word vectors are high-dimensional and dense vectors, calculating semantic similarity using Euclidean distance and others will be affected by the numerical magnitude, while cosine similarity only considers the direction and is not affected by the magnitude, so it can more stably measure the semantic similarity between words.
[0139] Through this method, the model can better understand the context semantics of the text, thereby generating high-quality summaries that are more consistent with the semantics of the article.
[0140] The present invention designs a module for enhancing semantic information, enabling the generator to better understand the context semantics of the text. This module draws on the method of TextRank algorithm for extracting keywords in extractive text summarization. Through the combination, these keywords extracted by the word embedding model can effectively reflect the core content and theme of the article. By combining this module, the TST-CycleGAN (TextSummarization-TextRank-CycleGAN) model is proposed. By combining this method with abstractive text summarization, the generator can more accurately grasp the main idea of the article and generate text summaries that are highly consistent with the theme.
[0141] The loss function of the TST-CycleGAN model adopts the cross-entropy loss function, as shown in formula (4):
[0142]
[0143] In formula (4), L(a) represents the cross-entropy loss, a represents a single text, and p i (a) represents the probability of the true class i, represents the probability of the model predicting class i.
[0144] As shown in Figures 4(a) and 4(b), the training process of the TST-CycleGAN model is as follows:
[0145] After inputting the article into the TST-CycleGAN model, first obtain the identity mapping loss L AA (a) through the generator F iM , and then pass through the generator G AS and the generator F SAGenerate fake articles cyclically to obtain the cycle consistency loss L NONPe (a);
[0146] Similarly, after inputting the abstract into the TST-CycleGAN model, L iM (s) and L NONPe (s) are obtained. Finally, the weighted sum with the adversarial loss L aMD of the GAN architecture is performed respectively to obtain:
[0147] Among them, L iM (s) and L iM (a) are the identity mapping losses of the abstract and the text respectively; L NONPe (s) and L NONPe (a) are the cycle consistency losses of the abstract and the text respectively; L aMD (F SA ) and L aMD (G AS ) are the adversarial losses of the two generators respectively.
[0149] In the TST-CycleGAN model, the adversarial loss L aMD is calculated through the adversarial game between the generator and the discriminator to ensure that the generated abstract is closer to the real abstract and improve the readability and authenticity of the generated text. Specifically, the discriminator D 1 is used to distinguish between real text and generated machine text, while the discriminator D 2 is used to distinguish between real abstract and generated machine abstract. During the training process, the generator G AS tries to generate high-quality abstracts so that the discriminator D 2 cannot distinguish its authenticity. At the same time, the generator F SA tries to generate text close to the original text so that the discriminator D 1 is difficult to distinguish. The adversarial loss uses the cross-entropy loss function to optimize the generator.
[0150] This loss acts on the TST-CycleGAN model through backpropagation, and can train the two generators separately, thereby improving the generation effect of the TST-CycleGAN model.
[0151] After inputting the article into the TST-CycleGAN model, a warm-up session is first carried out to enable the model to learn the rules of generating abstracts, and then large-scale unsupervised learning is carried out for training.
[0152] The warm-up session of supervised learning specifically includes:
[0153] The supervised training of the generator is mainly divided into two steps:
[0154] First step, input the text sequence of the real text and the real summary label into the generator G AS and F SA respectively. Through the forward propagation calculation of the BiLSTM encoding-decoding structure (Encoder-Decoder), the losses L AS and L SA output by the generator G Q and L Y are obtained;
[0155] The BiLSTM encoding-decoding structure is a sequence-to-sequence (Seq2Seq) framework based on the bidirectional long short-term memory network (BiLSTM), and is commonly used in text summarization, machine translation, and text generation tasks. This structure consists of a bidirectional LSTM encoder (BiLSTM Encoder) and an LSTM decoder (LSTM Decoder). The encoder processes the input text through forward and backward LSTMs, captures context information, and converts it into a fixed-length hidden state vector. The decoder then uses this hidden state vector to generate the target text. BiLSTM has a stronger information capture ability than unidirectional LSTM, can effectively learn long-distance dependencies, and improve the semantic integrity and coherence of summary generation. The present invention adopts an existing BiLSTM encoding-decoding framework and combines adversarial learning and cyclic consistency loss.
[0156] Second step, optimize the generator using the policy gradient method in reinforcement learning: input the real summary into the discriminator D 1 to obtain the prediction result, regard the prediction result as the reward signal, then calculate the logarithmic probability of this reward signal to calculate the policy gradient loss; finally, multiply the mean of the reward signal and the mean of the logarithmic probability to obtain the policy gradient loss. In the process of optimizing text summarization by reinforcement learning, the calculation of the policy gradient loss mainly depends on the reward signal and the logarithmic probability. First, the model generates a summary according to the current policy and calculates the reward signal of the summary using a certain evaluation criterion (such as ROUGE score or human scoring), that is, the score measuring the quality of the summary. This reward signal reflects the similarity between the summary generated by the model and the high-quality summary. The higher the score, the better the quality of the generated summary.
[0157] Next, calculate the logarithmic probability when the model generates the current summary, that is, the probability of the word predicted by the model at each step in the current context. This probability is usually calculated by the output layer of the model and measures the generation difficulty of the entire summary by accumulating the probabilities of all predicted words. The role of the logarithmic probability is to make the model more inclined to generate higher-quality summaries in future training processes.
[0158] Then, the loss is calculated by the policy gradient method, that is, the mean of the reward signal and the log probability are multiplied to form the final policy gradient loss. The goal of this loss is to adjust the model parameters so that summaries with high rewards are more likely to be generated, while the generation probability of low-quality summaries is gradually reduced.
[0159] Optimize the generator G AS and F SA as shown in equations (5) and (6) respectively:
[0160]
[0161] In equations (5) and (6), θ is the weight parameter; are the policy gradient losses for calculating the summary and the text respectively; L Q and L Y are the losses of the forward propagation output of the BiLSTM generator.
[0162] The training process of unsupervised learning specifically includes:
[0163] First, the original text is input into the module for enhancing semantic information to obtain the enhanced text. Then, the enhanced text and the original summary are respectively input into the tokenizer to obtain the text sequence a = (a 1 , a 2 ,..., a / ) and the summary sequence s = (s 1 , s 2 ,..., s / );
[0164] Then, a and s are respectively input into the generators G AS and F SA to generate the machine summary sequence s ′ = (s 1 ′ , s ′ 2 ,..., s ′ / ) and the machine text sequence a ′ = (a ′ 1 , a ′ 2 ,..., a ′ / ); To generate the authenticity of the summary, the cyclic consistency loss needs to be calculated.
[0165] Finally, the generated machine summary sequence s ′ and the machine text sequence a ′ are respectively input into F SA and G AS, generate fake text a similar to the original text and abstract fake and fake abstract s fake ;
[0166] After the original text and the original abstract have been transformed twice, the text and abstract generated again will inevitably be somewhat missing. To ensure the accuracy of the generation process, a cyclic consistency loss needs to be introduced to optimize the generator. Generally speaking, cyclic consistency means ensuring that the original text and the fake text after two transformations are almost identical. After the above training process, we finally obtain a fake and s fake After that, we perform processing on the loss function. We use the cross - entropy loss function to calculate the cyclic consistency loss L fake Ua, a NONPe_a V and the cyclic consistency loss L fake of the fake abstract s fake (s, s NONPe_` ); During the entire unsupervised training process, the two generators G fake and F AS and F SA are both involved in the generation process of the text and the abstract. Therefore, L NONPe_a and L NONPe_` need to optimize the two generators simultaneously.
[0167] In the training session of unsupervised learning, since there is no manually labeled training data, the generators G AS and F SA need to process not only text and abstract, but also face the situation of abstract input to G AS and text input to F SA . In this case, the input sequence faces the problem of over - modification by the generator, resulting in a decline in the quality of the generated text sequence. To solve this problem, an identity mapping loss is introduced, that is, when the text and the abstract are already "ideal", no further modification is required. It includes:
[0168] First, input the text sequence a=(a 1 , a 2 , …, a / ) and the abstract sequence s=(s 1 , s 2 , …, s / ) into the generators F SA and G AS respectively to obtain the output sequences a iM and s iM ;
[0169] Subsequently, calculate the identity mapping loss and respectively to optimize the generator F SAand G AS 。
[0170] The core of traditional generative adversarial network training is the adversarial loss, which is used to calculate the similarity between the generated fake text and fake summary and the original text and original summary. The generator G AS and F SA The generated fake summaries and fake texts are respectively input into the discriminators D 1 and D 2 to obtain the prediction results, and then calculate the loss with the true labels.
[0171] In summary, the loss function for optimizing the generator includes cyclic consistency loss, identity mapping loss, and adversarial loss, and the formula is as follows:
[0172]
[0173] where α and β are weight coefficients; L NONPe = ωL NONPe_a + ωL NONPe_` , ω is a weight coefficient; L NONPe , respectively refer to cyclic consistency loss, identity mapping loss of the summary, and adversarial loss of the discriminator D 1 backward optimized, is the identity mapping loss of the text, is the adversarial loss of the discriminator D 2 backward optimized, calculated through the cross-entropy loss function; refers to the loss function of the generator G AS , refers to the loss function of the generator F SA .
[0174] During the training process of the discriminator,
[0175] First, use the generator to generate fake summaries and fake texts, and use the Jieba tokenizer to tokenize the real text and real summary;
[0176] Then use the discriminator D 1 and D 2 composed of a three-layer convolutional neural network to generate prediction results, and this prediction result is the predicted value after softmax processing; the discriminators D 1 and D 2 adopt a three-layer convolutional neural network (CNN) structure, which are respectively used to judge whether the text and summary are real data. The core task of the discriminator is to distinguish artificial text / summary from the text / summary generated by the generator, so as to optimize the adversarial learning process of the model.
[0177] Each discriminator consists of three convolutional layers. Each convolutional layer uses multiple one-dimensional convolutional kernels to extract features from the text sequence and performs a non-linear transformation through the ReLU activation function to enhance the feature expression ability. The first convolutional layer is used to extract basic word-level features, the second convolutional layer captures phrase-level context relationships, and the third convolutional layer further extracts long-distance dependency information to ensure that the discriminator can effectively distinguish the authenticity of the text. After convolutional calculation, the feature vector is fed into the fully connected layer, and a probability value is output through the Sigmoid function to determine whether the input text is real text or generated text. During training, the discriminator is optimized through adversarial loss.
[0178] Next, set the real text and summary labels to 1, and the fake text and summary labels to 0, and calculate the loss L between the prediction results of the previous step and the true and false labels. g ; The formula is as follows:
[0179]
[0180] where γ is the weight coefficient; and are the losses between the predicted values of real and fake texts and the true and false labels respectively.
[0181] The loss calculation of the discriminator is mainly used to optimize its ability to distinguish real texts, summaries from generated texts, summaries. During training, first set the labels of real texts and real summaries to 1, indicating that they are from manually annotated data; at the same time, set the labels of machine-generated texts and summaries to 0, indicating that they are data generated by the generator. The discriminator needs to learn how to correctly classify these inputs, so it needs to calculate the loss between the prediction results and the true labels to optimize its judgment ability.
[0182] The loss of the discriminator consists of two parts: real data loss and generated data loss. First, the discriminator makes predictions on real texts and real summaries, outputs a probability value indicating the likelihood that they are "real", and calculates the error between this predicted value and the true label 1 to form the real data loss. Second, the discriminator makes predictions on the generated texts and summaries and calculates the error between their predicted values and the true label 0 to form the generated data loss. These two losses are weighted and summed according to a hyperparameter weight coefficient to obtain the final discriminator loss.
[0183] During training, the goal of the discriminator is to minimize this loss, that is, to improve the judgment accuracy of real data and better identify generated data. The generator is then continuously optimized accordingly, making it difficult for the discriminator to distinguish between machine-generated texts and real texts, thereby improving the quality of the generated summaries and making them more natural and fluent.
[0184] In the training stage of the CycleGAN framework of the present invention, the hardware environment used is NVIDIA RTX 4070 with 12GB of video memory, and the software environment is Python 3.8, Pytorch 2.0, and the LCSTS dataset. Initially, 500 pairs of text-summary pairs are selected as the warm-up part of the dataset, and 1000 pairs of text-summary pairs are used as the training dataset, where the validation set and test set data are both 10%. The following same hyperparameter settings are used during training: the batch_size is set to 4; the initial learning rate is 5×10-5, and the Adam optimizer is selected, which adaptively adjusts the learning rate during training until the loss functions of the generator and discriminator converge.
[0185] In the field of images, CycleGAN has achieved excellent results in unsupervised learning by virtue of its unique identity mapping loss and cycle consistency loss. However, due to the particularity of the text summarization task, relying solely on unsupervised learning may make it difficult to accurately capture the refined features of the summary, and the quality of the generated summaries often varies, unable to ensure the consistency of the summary with the original text in terms of relevance and conciseness. Therefore, the present invention adopts a semi-supervised learning method that combines a small amount of labeled data and a large amount of unlabeled data, which is particularly suitable for this framework.
[0186] The semi-supervised learning process of the present invention includes two parts:
[0187] (1) Supervised learning part: The present invention first uses a small part of paired data to warm up the model, compares the fake summary generated by the generator with the real summary, calculates the loss, and adjusts the model through backpropagation. Through this warm-up stage, the model can learn the conversion rules from text to summary, so that it can generate more accurate and meaningful summaries in the subsequent generation process, while significantly improving the coherence and accuracy of the summary.
[0188] (2) Unsupervised learning part: First, the original text is input into the enhanced semantic information module to generate an encoded text sequence. This sequence is encoded together with the original text, and the two encoded sequences are concatenated as the input of the generator. Through the generator G AS the machine summary s ′ can be obtained. Then, taking s ′ as the input, the fake article a SA is generated through the generator F fake . Similarly, after inputting the real summary into the model, the corresponding fake summary s fake can be obtained. Through the cycle consistency mechanism, the cycle consistency losses of L NONPe_ a Ua, a fake V and L NONPe_` (s, s fake ) are calculated respectively. In addition, the article and the summary respectively pass through FSA After that, the identity mapping loss can also be calculated to obtain AS Finally, the discriminator is responsible for distinguishing the generated fake abstracts and fake articles. Through the adversarial loss, combined with the cycle consistency loss and the identity mapping loss, the three losses are weighted and summed to obtain the final loss Finally, the discriminator is responsible for distinguishing the generated fake abstracts and fake articles. Through the adversarial loss, combined with the cycle consistency loss and the identity mapping loss, the three losses are weighted and summed to obtain the final loss with And the generator is trained through backpropagation. Finally, the model can generate meaningful abstracts.
[0189] In addition, the algorithm of the present invention can be applied as a core technology to a user question-answering system. The user only needs to input a large amount of text content to be interpreted, and the system can extract key information based on the algorithm of the present invention and generate a concise and refined abstract for the user to refer to, thereby helping the user to process various affairs more efficiently.
[0190] Embodiment 3
[0191] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of an abstractive text summarization method based on semi-supervised learning described in Embodiment 1 or 2 are implemented.
[0192] Embodiment 4
[0193] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of an abstractive text summarization method based on semi-supervised learning described in Embodiment 1 or 2 are implemented.
[0194] Embodiment 5
[0195] An abstractive text summarization system based on semi-supervised learning includes:
[0196] A data acquisition and preprocessing module, configured to: acquire data and perform preprocessing;
[0197] An automatic text summarization module, configured to: input the preprocessed data into the trained TST-CycleGAN model to implement automatic text summarization.
Claims
1. An abstract text summarization method based on semi-supervised learning, characterized in that: include: Acquire data and preprocess it; The preprocessed data is input into the trained TST-CycleGAN model to achieve automatic text summarization.
2. The method for abstract text summarization based on semi-supervised learning according to claim 1, characterized in that: In the TST-CycleGAN model training, the training dataset is the LCSTS dataset; Further preferably, the TST-CycleGAN model includes two pairs of generators and discriminators; one pair of generators and discriminators is used to convert articles into summaries, and the other pair of generators and discriminators is used to convert summaries into articles; In the generator-discriminator, the generator adopts the Seq2Seq structure. The generator includes an encoder and a decoder. The encoder adopts a bidirectional long short-term memory network, and the decoder adopts an LSTM with an attention mechanism. The encoder is used to capture the contextual dependency of the input sequence, and the decoder is used to generate a target summary based on the input sequence. The discriminator consists of a three-layer convolutional neural network to extract local features in the summary.
3. The method for abstract text summarization based on semi-supervised learning according to claim 2, characterized in that: The TST-CycleGAN model includes: a generator G for text generation summary AS , a generator F for summarizing text SA , a discriminator D1 for distinguishing true summaries from false summaries and a discriminator D2 for distinguishing true text from false text; in TST-CycleGAN model training, The first step is to warm up the supervised learning. The supervised learning is optimized through reinforcement learning. Paired text and summary sequences are input to enable the CycleGAN model to learn the mapping relationship between text and summary sequences and initially master the rules for generating summaries. The second step is to conduct unsupervised learning training, including: First, the original text a is input into the module for enhancing semantic information, and the keywords are extracted using the TextRank algorithm and concatenated with the original text to obtain text a′. Then, the text a′ and the original summary s sequence are input into two generators G respectively. AS and F SA , generate machine summaries and machine texts; Then, the original text a and the original summary s sequence, as well as the generated machine summary and machine text, are input into the discriminator D respectively. a and the discriminator D S , and calculate the adversarial loss; Then, the generated machine summary and machine text are input into F SA and G AS , generate a fake text a that restores the original text and the original summary fake and fake summaries fake ; Finally, the cycle consistency loss is used to calculate a and a respectively. fake , s and s fake The error between them, while using the identity mapping loss to constrain the generator, so that when the original text a is directly input to the generator F SA When the original summary s is directly input to the generator G AS , the output is s.
4. The method for abstract text summarization based on semi-supervised learning according to claim 2, characterized in that: The TST-CycleGAN model also includes a module that adds enhanced semantic information, namely the TextRank algorithm; First, after adding the module to enhance semantic information to the article input, text preprocessing operations are performed, including word segmentation and removal of stop words. The BERT model is used to extract the semantic information of the text sequence, and a vocabulary graph is constructed by calculating the semantic similarity, treating each word as a node in the graph. Then, the edges of the graph are established based on the co-occurrence relationship of the words; the co-occurrence relationship is based on the window size, and when two words appear in the window at the same time, an edge is established between the two words; Next, the initial score of each node is set to 1, and then the weight of each node is iteratively calculated using formula (1): In formula (1), S(V i ) represents the node V i score, d is the damping factor, In(V i ) and Out(V j ) respectively represent the nodes V i The connected input and output nodes, V i 、V j Represent two different nodes respectively; After iterative calculation, sort the scores of each node from large to small, and select the top n words as keywords; Finally, these keywords are encoded and concatenated with the text sequence to generate a new text sequence; Further preferably, the BERT model is used to extract semantic information of the text sequence, and a vocabulary graph is constructed by calculating semantic similarity; include: The text sequence x=(x1,x2,...,x T ) is input into the BERT model to obtain the contextual semantic representation of each word in the text sequence x, as shown in formula (2): h i =BERT(x i ) (2); In formula (2), x i represents the i-th word in the text, h i Represents x i High-dimensional word vectors; The semantic similarity between two nodes is calculated using cosine similarity, as shown in formula (3): In formula (3), and Represent two vectors respectively, and the numerator represents and The dot product of and The modulus length is normalized, and the value range of semantic similarity is between [-1,1]; Further preferably, the loss function of the TST-CycleGAN model adopts a cross entropy loss function, as shown in formula (4): In formula (4), L(a) represents the cross entropy loss, a represents a single text, and p i (a) represents the probability of the true category i, Represents the probability of the model predicting category i.
5. The method for abstract text summarization based on semi-supervised learning according to claim 1, characterized in that: The training process of the TST-CycleGAN model is as follows: After the article is input into the TST-CycleGAN model, it is first passed through the generator F SA Get the identity mapping loss L id (a), and then pass through the generator G AS Follow the generator F SA Generate fake articles in a loop and get the cycle consistency loss L cycle (a); Similarly, after inputting the summary into the TST-CycleGAN model, we get L id (s) with L cycle (s), and finally the adversarial loss L of the GAN architecture adv Performing weighted summation, we get: Among them, L id (s) and L id (a) Identity mapping loss for summary and text respectively; L cycle (s) and L cycle (a) Cycle consistency loss for summary and text respectively; L adv (F SA ) and L adv (G AS ) are the adversarial losses of the two generators respectively.
6. The method for abstract text summarization based on semi-supervised learning according to claim 1, characterized in that: The warm-up phase of supervised learning includes: The first step is to use the generator G AS and F SA The text sequence of the real text and the real summary label are input respectively, and the generator G is calculated through the forward propagation of the BiLSTM encoding-decoding structure. AS and F SA Output loss L Q and L F ; The second step is to optimize the generator using the policy gradient method in reinforcement learning: input the true summary into the discriminator D1 to obtain the prediction result, and use the prediction result as the reward signal, and then calculate the logarithmic probability of this reward signal to calculate the policy gradient loss; finally, multiply the mean of the reward signal and the mean of the logarithmic probability to obtain the policy gradient loss; For the generator G AS and F SA Optimize as shown in formula (5) and formula (6): In formula (5) and formula (6), θ is the weight parameter; They are the policy gradient losses for calculating summary and text respectively; L G and L F is the loss of the forward propagation output of the BiLSTM generator; Further preferably, the training process of unsupervised learning specifically includes: First, the original text is input into the module for enhancing semantic information to obtain enhanced text. Then, the enhanced text and the original summary are input into the word segmenter respectively to obtain the text sequence a = (a1, a2, ..., a T ) and the digest sequence s=(s1,s2,...,s T ); Then, a and s are input into the generator G respectively. AS and F SA , generate a machine summary sequence s′=(s′1,s′2,...,s′ T ) and the machine text sequence a′=(a′1,a′2,...,a′ T ); Finally, the generated machine summary sequence s′ and machine text sequence a′ are input into F SA and G AS , generate fake text a similar to the original text and summary fake and fake abstracts fake ; Process the loss function and use the cross entropy loss function to calculate the false text a fake The cycle consistency loss L cycle_a (a,a fake ) and fake abstracts fake The cycle consistency loss L cycle_s (s,s fake ); Further preferably, in the training phase of unsupervised learning, identity mapping loss is introduced, including: First, the text sequence a=(a1,a2,…,a T ) and the digest sequence s=(s1,s2,…,s T ) respectively input the generator F SA and G AS Get the output sequence a id and id ; Then, the identity mapping loss is calculated as well as Optimize the generator F separately SA and G AS ; The loss function for optimizing the generator includes cycle consistency loss, identity mapping loss, and adversarial loss, and the formula is as follows: Among them, α and β are weight coefficients; L cycle =ωL cycle_a +ωL cycle_s , ω is the weight coefficient; L cycle , They refer to the cycle consistency loss, the identity mapping loss of the summary, and the adversarial loss of the discriminator D1 reverse optimization, respectively. is the identity mapping loss for text, It is the adversarial loss of the reverse optimization of the discriminator D2, calculated by the cross entropy loss function; refers to the generator G AS The loss function is refers to the generator F SA The loss function of .
7. The method for abstract text summarization based on semi-supervised learning according to claim 2, characterized in that: During the training of the discriminator, Firstly, the generator is used to generate fake summaries and fake texts, and Jieba word segmenter is used to segment the real texts and real summaries. Then, the discriminators D1 and D2 composed of three-layer convolutional neural networks are used to generate prediction results, which are the prediction values after softmax processing; Next, set the real text and summary labels to 1, the fake text and summary labels to 0, and calculate the loss L between the prediction results of the previous step and the real and fake labels D ; The formula is as follows: Among them, γ is the weight coefficient; and are the losses between true and false text prediction values and true and false labels, respectively.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of an abstract text summarization method based on semi-supervised learning described in any one of claims 1-7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of an abstract text summarization method based on semi-supervised learning described in any one of claims 1 to 7 are implemented.
10. An abstract text summarization system based on semi-supervised learning, characterized in that: include: The data acquisition and preprocessing module is configured to: acquire data and perform preprocessing; The automatic text summarization module is configured to: input the preprocessed data into the trained TST-CycleGAN model to achieve automatic text summarization.