Automatic Extraction Method for Text Summarization Based on Dual-Stream Attention and Positional Residual Connection
By introducing the RBPSum model with dual-flow attention and position residual connection, the problem that the existing extracted text summary model cannot effectively utilize sentence variability and previous selected information is solved, and high-accurate text summary extraction on small-scale data sets is achieved.
Patent Information
- Application Number
- CN202210950607.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-09
AI Technical Summary
When extracting text key information, the existing extracted text summary model cannot effectively utilize the differences between sentences and previously selected sentence information, resulting in inaccurate extraction results.
The RBPSum model based on dual-flow attention and position residual connection is adopted, and the interactive information between sentences is captured through the dual-flow self-attention mechanism, and the position information is injected through the position residual connection, combining the new objective function to optimize the model training process.
Under the condition of limited training data volume, the robustness of the model and the accuracy of extracting information are improved, especially on small-scale data sets, and the performance is better than that of existing models.
Smart Images

Figure CN115309887B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to an automatic text summarization extraction method based on dual-stream attention and position residual connection. Background Art
[0002] In the field of Natural Language Processing (NLP), the goal of the automatic text summarization task is to analyze the overall semantics of a document and convert a long text into a summary containing the key information of the original document. The application scenarios of text summarization are very broad, such as automatically generating news headlines, literature reports, e-commerce marketing introduction content, etc.
[0003] In recent years, how to generate high-quality text summaries has been widely studied by researchers. Currently, there are mainly two mainstream implementation methods: extractive text summarization and generative text summarization.
[0004] Generative text summarization generates summaries word by word in an autoregressive manner of a language model. The generative summary method can be regarded as a reconstruction process, which is mainly based on the Sequence-to-Sequence model. The first step is to use an encoder to encode the semantic information of the article, and then use a decoder to generate a summary according to the semantic information. The generative method has high continuity and freedom (the decoder can generate new words). However, this method requires the model to have strong natural language understanding and natural language generation capabilities, and this method also often suffers from problems such as inconsistency, resulting in the generation of incorrect information.
[0005] Extractive text summarization generates a summary by selecting key content with representative information in the article, such as whole sentences and clauses, and splicing and combining these information. Extractive summarization can be defined as a sentence ranking problem. It encodes and extracts the semantics of the original text, then calculates the key scores of each sentence through semantic features, and finally selects the k sentences with the highest scores as the summary. For the extractive method, since the sentences are directly extracted from the original text, it has better coherence and readability. However, in order to obtain sentence representations, most current extractive models rely on the Transformer encoder to extract the feature representations of sentences. Although the model structure of the Transformer encoder endows it with parallel computing capabilities, the model has no ability to distinguish the differences between sentences when processing sentences. That is, when calculating the representation vector of each sentence, the vector representations of all sentences are calculated at once through parallel computing, and the information of the previously selected sentences cannot be taken into account, which makes some information lose its reference value and makes the final extraction result inaccurate. Summary of the Invention
[0006] In view of the above problems in the prior art, the technical problem to be solved by the present invention is: how to more accurately extract key information in the text as the text summary.
[0007] To solve the above technical problem, the present invention adopts the following technical solutions:
[0008] An automatic text summary extraction method based on dual-stream attention and position residual connection, comprising the following steps:
[0009] S100: Select a publicly available dataset, which includes D texts and the corresponding actual summary information for each text; each text contains several sentences, and all sentences included in D are marked with original labels; randomly select a part from the dataset with original labels as the training set, and the remaining part as the test set;
[0010] S200: Construct the RBPSum model, which includes a sentence encoder, a context encoder, and an output layer;
[0011] The context encoder consists of L sentence enhancement layers, and each sentence enhancement layer consists of multiple Transformer encoders. The attention mechanism used in the Transformer encoder is dual-stream self-attention;
[0012] The output layer includes a position residual connection module and a probability prediction module;
[0013] S300: Assume that the training set contains P texts and the training data batch is Q. Divide P into Q equal parts to get A, that is, A = P / Q. Initialize the RBPSum model:
[0014] S400: Let batch = 1;
[0015] S410: Select A training samples from the training set as a batch, batch ∈ [1, Q];
[0016] S420: Let t = 1;
[0017] S430: Select the t-th text D from the training set t , and use the sentence encoder to extract features from all sentences in D t to obtain the sentence feature representation E of each sentence in D t , where n represents the number of sentences contained in all texts in D t :n , t ;
[0018] S440: Use E t :n as the input of the context encoder, Et :n Pass through L sentence enhancement layers in sequence and output to obtain E t :n Context feature encoding information at the document level;
[0019] S450: Use the output layer to calculate and output D t The probability values P of all statements in
[0020] S451: Use the position residual connection module to add E t :n to its corresponding context feature encoding information to obtain the sentence context vector after position enhancement corresponding to E t :n The calculation expression is as follows: The calculation expression is as follows:
[0021]
[0022] where β is a hyperparameter and X is the output vector of the sentence enhancement layer;
[0023] S452: Use the probability prediction module to calculate the probability value P that each statement in D t is selected as the summary, and the expression is as follows:
[0024]
[0025] where σ represents the Sigmoid function and P ∈ (0, 1);
[0026] S453: Sort the P values of all statements in D t in descending order, and select the statements corresponding to the top K probability values as the predicted summary of D t ;
[0027] S460: Construct the objective function object, and the specific steps are as follows:
[0028] S461: Calculate the classification loss function BCELoss between the predicted summary of D t and the corresponding actual summary;
[0029] S462: Use the similarity function to calculate the distance Dist1 between the actual summary of D t and the corresponding text feature representation, and the calculation expression is as follows:
[0030] Dist1 = CosSim(E tgt , E doc ); (3)
[0031] where E tgtDenote the actual abstract as E doc Denote the text feature representation, and CosSim(·) denote the cosine similarity function;
[0032]
[0033] where n is the number of sentences in the text, q = (1, …, n), and E q Denote D t The text feature representation of the q-th sentence in the text;
[0034] S463: Calculate the distance Dist2 between the sentences included in the actual abstract of D t The calculation expression is as follows:
[0035]
[0036] where C i Denote the sentence context feature representation after position residual connection, m denotes the number of sentences in each actual abstract, i = (0, …, m - 1), j = i + 1, j = (1, …, m);
[0037] S464: Calculate the objective function of D t The expression is as follows:
[0038] object = BCELoss - Dist1 + Dist2; (6)
[0039] S465: Update the RBPSum model parameters by backpropagation according to the objective function;
[0040] S466: If t ≥ A, record the model parameters of the current RBPSum as a checkpoint for saving and execute the next step; if t < A, set t = t + 1 and return to S430;
[0041] S470: If batch ≥ Q, execute the next step; if batch < Q, set batch = batch + 1 and return to S410;
[0042] S500: Determine the final parameters of the RBPSum model:
[0043] S510: Select one from the Q checkpoints as the parameters used by the current RBPSum model;
[0044] S520: Assume that there are W texts in the test set, select the s-th text from W as the input of the current RBPSum model, and output the predicted abstract of the s-th text;
[0045] S530: Calculate the ROUGE scores between each sentence in the predicted summary of the s-th text and the actual summary corresponding to the s-th text using the greedy algorithm, obtaining a number of ROUGE scores, and then calculate the arithmetic mean Range of these ROUGE scores;
[0046] S540: Calculate the Range corresponding to the W texts using the methods described in S520 and S530, and obtain the Range' corresponding to this checkpoint by calculating the arithmetic mean of the Range values corresponding to the W texts;
[0047] S550: Repeat S520 - S540, traverse all checkpoints, and calculate Q Range' values;
[0048] S560: Sort the Q Range' values in descending order, select the checkpoint corresponding to the highest Range' as the parameter of the current RBPSum model, and use this current RBPSum model as the finally trained RBPSum model;
[0049] S600: Input the text to be predicted into the finally trained RBPSum model, and the output is the predicted text summary of this text to be predicted.
[0050] Preferably, the sentence encoder used in S430 is the pre-trained RoBERTa.
[0051] The pre-trained RoBERTa is a model verified by a large amount of data. The results calculated by this model can be directly used without further processing of the initial data, saving a large amount of computing.
[0052] Preferably, the sequence of sentences segmented by the sentence encoder in S430 is [CLS]S1[SEP][CLS]S2,..., [SEP][CLS]S x [SEP], where S x represents the x-th sentence, and the output vectors at all [CLS] positions in this sequence are used as the feature representation E of each sentence :n .
[0053] RBPSum uses the pre-trained RoBERTa as the sentence encoder. In RoBERTa, the [CLS] token is used to gather the semantic information of each word or word segment in the sentence to serve as the representation vector of the sentence, and the [SEP] token is used to mark the end position of the sentence to segment different sentences.
[0054] Preferably, the specific steps to obtain the output vector X of the sentence enhancement layer in S451 are as follows:
[0055] S451-1: Add positional encoding to each sentence feature representation, with the expression as follows:
[0056]
[0057] where E :n = [E0, E1, E2,.., E n , and PosEnc represents the positional encoding function;
[0058] S451-2: Calculate the output vector X of the sentence enhancement layer, with the specific expression as follows:
[0059]
[0060]
[0061] where LN represents the layer normalization function, and BiStreamAttn represents the bi-stream self-attention module.
[0062] Compared with the prior art, the present invention has at least the following advantages:
[0063] 1. The present invention can ensure strong robustness of the model under the condition of limited training data volume. The bi-stream attention mechanism and the proposed objective function used in the present invention enable the RBPSum model to fully learn the information of the selected summary sentences in the previous time steps and the relationship information between key sentences and non-key sentences, thus further improving the accuracy of the RBPSum model; the positional residual connection ensures that under the condition of insufficient training data volume, more positional information is used to guide the RBPSum model to extract summaries by injecting additional positional information into the RBPSum model, ensuring good robustness of the RBPSum model.
[0064] 2. The present invention proposes a generation method RBPSum model for extractive text summarization. The performance of RBPSum is better than previous SOTA models on the CNN / DailyMail public dataset; at the same time, RBPSum can achieve good performance with only 1 / 3 of the training set and is better than most baseline models trained with the entire training set.
[0065] 3. The present invention proposes three methods, namely the bi-stream attention mechanism, the positional residual connection, and a new objective function, to further improve the model performance. Ablation experiments show that the methods we proposed can further improve the model performance, especially on small-scale datasets; at the same time, the method of the present invention can be conveniently migrated to other extractive summary generation models.
[0066] 4. When there is only a small-scale data set, the performance of the model of the present invention also has more accurate information extraction capabilities than the existing models. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is the RBPSum algorithm framework diagram. DETAILED DESCRIPTION
[0068] The present invention is described in further detail below.
[0069] At a certain time step, when deciding whether a sentence should be selected as a summary, the sentences that have been selected in the previous time step still need to be used as a reference, which is similar to the autoregressive method. To achieve this function, the present invention proposes a dual-stream attention mechanism, which is different from the previous implementation method directly using an autoregressive decoder; most extractive text summarization methods only use classification loss as the objective function for model training. The present invention takes into account that the overall semantics of a document should be similar to the semantics of its summary, and that the sentences in the summary should contain different key information in the document. Therefore, a new objective function is proposed, which can narrow the distance between the document semantic information and the summary semantic information on the basis of the classification loss, and widen the distance between the sentences in the summary, so as to improve the accuracy of extracting information.
[0070] See also Figure 1 , a text summarization automatic extraction method based on dual-stream attention and position residual connection, comprising the following steps:
[0071] S100: Select a public data set, which includes D texts and actual summary information corresponding to each text, where the actual summary of each text is a sentence annotated manually; each text includes several sentences, and all sentences included in D are marked with original labels; randomly extract a part from the data set with original labels as a training set, and the remaining part as a test set;
[0072] S200: Build the RBPSum model, which includes a sentence encoder, a context encoder, and an output layer;
[0073] The context encoder consists of L sentence enhancement layers, and each sentence enhancement layer consists of multiple Transformer encoders. The attention mechanism used in the Transformer encoder is the two-stream self-attention; the two-stream self-attention is a prior art. Since the sentence representation output by RoBERTa mainly focuses on word-level information, the interaction information between sentences is not rich. Therefore, a sentence enhancement layer is designed to capture the interaction information between sentences; the sentence enhancement layer consists of multiple Transformer encoders. Different from the traditional Transformer encoder, in the present invention, the multi-head self-attention module of the Transformer is modified and extended to a two-stream self-attention module.
[0074] The output layer includes a position residual connection module and a probability prediction module;
[0075] S300: Assume that the training set contains P texts and the training data batches are Q. Divide P equally into Q to get A, that is, A = P / Q, and initialize the RBPSum model:
[0076] S400: Let batch = 1;
[0077] S410: Select A training samples from the training set as a batch, where batch ∈ [1, Q];
[0078] S420: Let t = 1;
[0079] S430: Select the t-th text D from the training set t , and use the sentence encoder to extract features from all statements in D t to obtain the sentence feature representation E of each statement in D t , where n represents the number of sentences included in all texts in D t :n , where n represents the number of sentences included in all texts in D t ;
[0080] The sentence encoder used in S430 is the pre-trained RoBERTa, and the pre-trained RoBERTa is a prior art.
[0081] In S430, the sentence encoder divides the sentence sequence into [CLS]S1[ESP][CLS]S2,..., [ESP][CLS]S x [ESP], where S x represents the x-th statement, and the output vectors at all [CLS] positions in this sequence are used as the feature representation E of each sentence :n. RBPSum uses the pre-trained RoBERTa as the sentence encoder. In RoBERTa, the [CLS] token is used to pool the semantic information of each word or sub-word in the sentence to serve as the sentence representation vector, and the [SEP] token is used to mark the end position of the sentence to segment different sentences.
[0082] S440: Take E t :n as the input of the context encoder. E t :n Passes through L sentence enhancement layers in sequence, and the output is E t :n The context feature encoding information at the document level. The sentence enhancement layer is spliced on top of the sentence encoder. The sentence enhancement layer only focuses on the sentence representation, that is, the vector at the position of the [CLS] identifier in the output vector of the sentence encoder. Therefore, it can capture the context information of bidirectional sentences. At the same time, the introduction of two-stream self-attention enables it to capture the information of the sentences selected at the previous moment;
[0083] S450: Use the output layer to calculate and output the probability values P of all statements in D t as follows:
[0084] S451: Use the position residual connection module to perform a summation operation on E t :n and its corresponding context feature encoding information to obtain the sentence context vector after position enhancement corresponding to E t :n The calculation expression is as follows: The calculation expression is as follows:
[0085]
[0086] where β is a hyperparameter and X is the output vector of the sentence enhancement layer;
[0087] The specific steps to obtain the output vector X of the sentence enhancement layer in S451 are as follows:
[0088] S451-1: Add position encoding to each sentence feature representation. The expression is as follows:
[0089]
[0090] where, E :n =[E0, E1, E2,.., E n , PosEnc represents the position encoding function; adding position encoding to each sentence feature representation is used to distinguish the positional relationship between each sentence at the document level.
[0091] S451 - 2: Calculate the output vector X of the sentence enhancement layer. The specific expression is as follows:
[0092]
[0093]
[0094] Among them, LN represents the layer normalization function, and BiStreamAttn represents the bi - stream self - attention module.
[0095] S452: Use the probability prediction module to calculate P, the probability value that each statement in D t is selected as the summary, and the expression is as follows:
[0096]
[0097] Among them, σ represents the Sigmoid function, and P ∈ (0, 1);
[0098] S453: Sort the P values of all statements in D t in descending order, and select the statements corresponding to the top K probability values as the predicted summary of D t . Generally, the value of K is taken as 3;
[0099] S460: Construct the objective function object. The specific steps are as follows:
[0100] S461: Calculate the classification loss function BCELoss between the predicted summary of D t and the corresponding actual summary. BCELoss represents Binary Classification Entropy, which is a prior art;
[0101] S462: Use the similarity function to calculate the distance Dist1 between the actual summary of D t and the corresponding text feature representation. The similarity function is a prior art, and the calculation expression is as follows:
[0102] Dist1 = CosSim(E tgt , E doc ); (3)
[0103] Among them, E tgt represents the actual summary, E doc represents the text feature representation, and CosSim(·) represents the cosine similarity function;
[0104]
[0105] Among them, n is the number of sentences in the text, q = (1,..., n), and E q represents D tThe text feature representation of the q-th sentence in the text;
[0106] S463: Calculate D t the distance Dist2 between the sentences included in the actual abstract of D, and the calculation expression is as follows:
[0107]
[0108] where C i represents the sentence context feature representation after position residual connection, m represents the number of sentences in each actual abstract, i = (0,..., m - 1), j = i + 1, j = (1,..., m);
[0109] S464: Calculate the objective function of D t The expression is as follows:
[0110] object = BCELoss - Dist1 + Dist2; (6)
[0111] S465: Update the RBPSum model parameters by backpropagation according to the objective function;
[0112] S466: If t ≥ A, record the model parameters of the current RBPSum as a checkpoint for saving and execute the next step; if t < A, set t = t + 1 and return to S430;
[0113] S470: If batch ≥ Q, execute the next step; if batch < Q, set batch = batch + 1 and return to S410;
[0114] S500: Determine the final parameters of the RBPSum model:
[0115] S510: Select one from the Q checkpoints as the parameters used by the current RBPSum model;
[0116] S520: Assume that there are W texts in the test set, select the s-th text from W as the input of the current RBPSum model, and output the predicted abstract of the s-th text;
[0117] S530: Use the greedy algorithm to calculate the ROUGE score between each sentence in the predicted abstract of the s-th text and the actual abstract corresponding to the s-th text, obtain several ROUGE scores, and then calculate the arithmetic mean Range of these several ROUGE scores;
[0118] S540: Calculate the Range corresponding to W texts using the methods described in S520 and S530, and obtain the Range' corresponding to this checkpoint by taking the arithmetic mean of the Ranges corresponding to the W texts;
[0119] S550: Repeat S520 - S540 to traverse all checkpoints and calculate Q Range';
[0120] S560: Sort the Q Range' in descending order, select the checkpoint corresponding to the highest Range' as the parameter of the current RBPSum model, and use this current RBPSum model as the finally trained RBPSum model;
[0121] S600: Input the text to be predicted into the finally trained RBPSum model, and the output is the predicted text summary of the text to be predicted.
[0122] Data Experimental Analysis
[0123] 1. Dataset
[0124] The unannotated version of CNN / DailyMail is selected as the dataset. The data is divided according to the method proposed by Hermann et al. The numbers of the training set / validation set / test set are 287,083 / 13,367 / 11,490 respectively; for each text document, sentence segmentation is performed through the Stanford CoreNLP toolkit; finally, the ROUGE F1 (R-1 / R-2 / R-L) scores between each sentence in the text and the reference summary are calculated using the greedy algorithm, and the 3 sentences with the highest scores are selected as the Oracle.
[0125] 2. Baseline Models
[0126] The following baseline models (Baselines) are selected for the experiment:
[0127] · Lead-3 is a baseline model for extractive text summarization, which selects the first 3 sentences in the text as the summary.
[0128] · Oracle is used as the upper bound of extractive text summarization. The ROUGE scores between each sentence in the text and the reference summary are calculated using the greedy algorithm, and the 3 sentences with the highest scores are selected as the summary.
[0129] ·RoBERTa+Transformer-Encoder is the contrast model proposed in this invention, which consists of RoBERTa and two layers of standard Transformer-encoders. This model maintains the same parameter settings as RBPSum, does not use a two-stream attention module or position residual connection, and only uses Binary Classification Entropy as the objective function.
[0130] ·REFRESH is an extractive text summarization model trained by globally optimizing the ROUGE score and combining reinforcement learning.
[0131] ·Bottom-UP is a generative text summarization model built on an encoder-decoder structure and uses Bottom-UP Copy attention to select key sentences.
[0132] ·BertSumExt is an extractive text summarization model based on the pre-trained model BERT concatenated with a Transformer-encoder.
[0133] ·HiStruct+RoBERTa-base is an extractive text summarization model based on RoBERTa, which enables the model to learn structural knowledge of the text by adding hierarchical information of the text.
[0134] 3. Training Details
[0135] The model is built using the Pytorch framework and RoBERTa is initialized with the pre-trained parameters of roberta-base. The dataset is tokenized using RoBERTa's tokenizer and the maximum length of the input sequence is limited to 514. The Adam optimizer (β1 = 0.9, β2 = 0.999) is used, and the learning rate setting refers to Vaswani et al. (warmup = 32000):
[0136] lr = 2e -3 ·min(step -0.5 , step·warmup -1.5 )(10)
[0137] The model was trained for 200,000 steps on a single GPU (GTX 3060 11G) with a gradient accumulation of 10 and a Batch Size of 4. During the training phase, the model saved a checkpoint every 5,000 steps. During the testing phase, the ROUGE F1 value was used as the evaluation metric, and all checkpoints were tested on the test set, and the checkpoint with the highest score was selected as the final model parameters. Meanwhile, during the inference phase, Trigram Blocking was performed to reduce sentence redundancy.
[0138] Since the sentence labels are not visible during the testing phase, only Mask2 was used during the testing phase, and the sentence enhancement layer at this time is equivalent to a standard Transformer encoder. For the hyperparameters α and β, α = 0.05, 0.15, 0.25, 0.5, 1.0 and β = 0.15, 0.25, 0.35, 0.45 were set, and it was found that the model performance was optimal when α = 0.05 and β = 0.15.
[0139] Table 1. ROUGE F1 results on the CNN / DailyMail test set
[0140]
[0141] 4. Hyperparameter Experiment
[0142] Some previous research works believed that the position information of sentences is very important for the extractive summarization task. In order to explore the impact of sentence position information on the task, a hyperparameter experiment was conducted. Specifically, the present invention respectively selected about 1 / 3 (90,003 documents), 1 / 5 (60,004 documents) and 1 / 9 (30,004 documents) of the data from the CNN / DailyMail training set, and set the hyperparameter β = 0.15, 0.25, 0.35, 0.45, and studied the impact of sentence position information on the model performance of different dataset sizes by controlling variables. To avoid model overfitting, the model training steps were set to 50,000 and warmup = 8,000 at this time.
[0143] Table 2. ROUGE F1 test results using 90,003 training data
[0144]
[0145] Table 3. ROUGE F1 test results using 60,004 training data
[0146]
[0147]
[0148] Table 4. ROUGE F1 test results using 30,004 training data
[0149]
[0150] 5. Experimental results
[0151] Since BertSumExt is trained with different batch sizes, gradient accumulation, and training steps, it cannot be directly compared with it. Therefore, this experiment retrains BertSumExt with the same parameter settings as RBPSum.
[0152] Table 1 shows the results of the model on the CNN / DailyMail test set. When using all the training sets, compared with BertSumExt, HiBERT, and HiStruct+RoBERTa-base, RBPSum improves R-1 / 2 / L by 0.53 / 0.39 / 0.49, 0.59 / 0.17 / 0.37, and 0.13 / 0.09 / 0.06 respectively. Notably, when using only 1 / 3 of the training set, compared with REFRESH, PGN, and Bottom-UP, RBPSum improves R-1 / 2 / L by 3.49 / 2.18 / 3.24, 3.96 / 3.10 / 3.46, and 2.27 / 1.70 / 1.50 respectively.
[0153] It can be seen that compared with the previous state-of-the-art extraction and generation methods, BPRsum can achieve better ROUGE F1 results (43.78 / 20.63 / 40.09 on R-1 / 2 / L respectively). Especially compared with most strong baseline models, PRSum can train better model performance with only about 1 / 3 of the CNN / DailyMail training set. On the other hand, the method proposed in the present invention further improves the performance of BertSumExt.
[0154] As can be seen from Table 2, Table 3, and Table 4, RBPSum achieved the best results: 1) trained on 30,004 documents with β = 0.45; 2) trained on 60,004 documents with β = 0.35; 3) trained on 90,003 documents with β = 0.25. It can be found that as the number of datasets decreases, the larger the β value, the better the model performance, which indicates that in the case of small-scale data, the model is more dependent on the position information of sentences.
[0155] In the case of limited dataset size, RBPSum shows a significant performance improvement compared to RoBERTa+Transformer-encode and BertSumExt. Compared to BertSumExt, RBPSum increases R-1 / 2 / L by 0.66 / 0.59 / 0.66 (in the case of 90,003 documents), 0.78 / 0.61 / 0.74 (in the case of 60,004 documents), and 0.76 / 0.66 / 0.69 (in the case of 30,004 documents).
[0156] Table 5. Ablation experiment results of RBPSum on the CNN / DailMail dataset
[0157] Model R-1 R-2 R-L RBPSum <![CDATA 43.78 > <![CDATA 20.63 > <![CDATA 40.09 > Only with our objective function 43.72 20.59 40.04 Only with Bi-Stream Attention 43.74 20.59 40.04 Only with Position Residual Connection 43.77 20.60 40.07 RoBERTa+Transformer-encoder 43.70 20.57 40.03
[0158] Table 6. Ablation experiment results of BertSUM on the CNN / DailMail dataset
[0159]
[0160] 6. Ablation experiment:
[0161] To fully demonstrate the effectiveness of the method proposed in the present invention, including the objective function, Bi-Stream Attention mechanism, and Position Residual Connection, ablation experiments were conducted on them based on the method of controlling variables.
[0162] The results are shown in Table 5: It can be seen that all the methods proposed in the present invention improve the performance of the model to varying degrees, especially the contribution of Position Residual Connection is the most significant; at the same time, to further prove the transferability of these methods, the present invention applied these methods to BertSumExt and observed the experimental results.
[0163] Table 6 summarizes the results of BertSUMExt, where the contribution of the objective function is the most significant.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An automatic text summarization extraction method based on dual-stream attention and positional residual connection, characterized in that: It includes the following steps: S100: Select a publicly available dataset, which includes D texts and the corresponding actual summary information for each text; each of the texts contains several sentences, and all the sentences included in D are labeled with original labels; randomly select a part from the dataset with original labels as the training set, and the remaining part as the test set; S200: Construct the RBPSum model, which includes a sentence encoder, a context encoder, and an output layer; The context encoder consists of L sentence reinforcement layers, and each sentence reinforcement layer consists of multiple Transformer encoders. The attention mechanism used in the Transformer encoder is the two-stream self-attention; The output layer includes a position residual connection module and a probability prediction module; S300: Assume that the training set contains P texts and the training data batch is Q. Divide P into Q equal parts to get A, that is, A = P / Q. Initialize the RBPSum model: S400: Let batch = 1; S410: Select A training samples from the training set as a batch, batch ∈ [1, Q]; S420: Let t = 1; S430: Select the t-th text D from the training set t , and use the sentence encoder to extract features for all sentences in D t to obtain the sentence feature representation E of each sentence in D t , where n represents the number of sentences contained in all texts in D t :n , and t ; S440: Use E t :n as the input to the context encoder, and E t :n successively passes through L sentence enhancement layers to output E t :n context feature encoding information at the document level; S450: Calculate and output D using the output layer t for all statements in S451: Use the position residual connection module to sum E t :n with its corresponding context feature encoding information to obtain the sentence context vector after position enhancement corresponding to E t :n The calculation expression is as follows: The calculation expression is as follows: Among them, β is a hyperparameter, and X is the output vector of the sentence reinforcement layer; S452: Calculate D using the probability prediction module t The probability value P that each statement in Among them, σ represents the Sigmoid function, P ∈ (0, 1); S453: Sort the P-values of all statements in D t in descending order, and select the statements corresponding to the top K probability values as the prediction summary for D t ; S460: Construct the objective function object, and the specific steps are as follows: S461: Calculate the classification loss function BCELoss between the predicted summary and the corresponding actual summary of D t ; S462: Calculate D using the similarity function t The distance Dist1 between the actual abstract of Dist1 = CosSim(E tgt , E doc ); (3) Among them, E tgt represents the actual abstract, and E doc represents the text feature representation, and CosSim(·) represents the cosine similarity function; where n is the number of sentences in the text, q = (1,..., n), and E q represents D t is the text feature representation of the q-th sentence in the text; S463: Calculate D t Calculate the distance Dist2 between the statements included in the actual abstract of Among them, C i represents the sentence context feature representation after positional residual connection, m represents the number of sentences in each actual abstract, i = (0, …, m - 1), j = i + 1, j = (1, …, m); S464: Calculate the objective function of D t as follows: object = BCELoss - Dist1 + Dist2; (6) S465: Update the RBPSum model parameters by backpropagation according to the objective function; S466: If t ≥ A, record the current model parameters of RBPSum as a checkpoint for saving and execute the next step; if t < A, let t = t + 1 and return to S430; S470: If batch ≥ Q, execute the next step; if batch < Q, let batch = batch + 1 and return to S410; S500: Determine the final parameters of the RBPSum model: S510: Select one from Q checkpoints as the parameters used by the current RBPSum model; S520: Assume that the test set contains W texts. Select the s-th text from W as the input of the current RBPSum model, and output the predicted summary of the s-th text; S530: Use the greedy algorithm to calculate the ROUGE score between each sentence in the predicted summary of the s-th text and the actual summary corresponding to the s-th text, obtain several ROUGE scores, and then calculate the arithmetic mean Range of these several ROUGE scores; S540: Calculate the Range corresponding to the W texts using the method described in S520 and S530, and calculate the arithmetic mean of the Range corresponding to the W texts to get the Range' corresponding to this checkpoint; S550: Repeat S520 - S540, traverse all checkpoints, and calculate Q Range'; S560: Sort the Q Range's in descending order, select the checkpoint corresponding to the highest Range' as the parameter of the current RBPSum model, and use this current RBPSum model as the finally trained RBPSum model; S600: Input the text to be predicted into the finally trained RBPSum model, and the output is the predicted text summary of the text to be predicted.
2. The automatic extraction method for text summarization based on dual-stream attention and position residual connection according to claim 1, characterized in that: In S430, the sentence encoder uses the pre-trained RoBERTa.
3. The automatic text summary extraction method based on dual-stream attention and position residual connection according to claim 2, characterized in that: In S430, the sentence encoder divides the sequence of sentences into [CLS]S1[ESP][CLS]S2, …, [ESP][CLS]S x [ESP], where S x represents the x-th statement, and the output vectors at all [CLS] positions in this sequence are used as the feature representation E of each sentence :n .
4. The automatic text summary extraction method based on dual-stream attention and position residual connection as described in claim 3, wherein: The specific steps for obtaining the output vector X of the sentence enhancement layer in S451 are as follows: S451-1: Add position encoding to each sentence feature representation, and the expression is as follows: Among them, E :n = [E0, E1, E2,.., E n , and PosEnc represents the position encoding function; S451-2: Calculate the output vector X of the sentence enhancement layer, and the specific expression is as follows: Among them, LN represents the layer normalization function, and BiStreamAttn represents the bi-stream self-attention module.
Citation Information
Patent Citations
Method for improving dialogue text generation based on text abstract generation and bidirectional corpus
CN113158665A
Two-stage hybrid automatic abstracting method for judicial judgment documents
CN114169312A