Sentence-level automatic summarization model system and summary generation method based on deep learning
By designing a clause-level automatic summary model system based on deep learning, using BERT and Transformer encoders and BERTScore matchers, the sentence importance scoring and redundancy problems in the prior art are solved, and high integrity and efficient summary generation are achieved.
Patent Information
- Application Number
- CN202210589278.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-26
AI Technical Summary
The existing extracted abstract model has room for improvement in processing the importance scores and redundant information of sentences in text, and the entire sentence-level extraction is easily trapped in the local optimality and cannot obtain the best summary.
A clause-level automatic digest model system based on deep learning is designed, and the original data is split into clauses through sentence-level extraction units. The vector representation of clauses is obtained using an encoder based on BERT and Transformer. Combining a classifier and a summary matcher based on BERTScore, a semantic best match summary is generated.
The redundancy problem was effectively solved, and the integrity and extraction speed of summary information were improved. The experimental results showed that the model was better than the comparison model in both automatic evaluation and manual evaluation, and the Rouge index was improved.
Smart Images

Figure CN115033659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a clause-level automatic summary model system and a summary generation method based on deep learning. Background Art
[0002] Automatic text summarization technology uses computers to summarize and conclude the given source text information and generate concise, fluent short texts that retain key information. It has become an important task in natural language processing today. It is mainly divided into two categories: extractive summarization and generative summarization. Generative summarization deeply analyzes the semantic information of the text and allows the generation of new words and phrases. Extractive summarization scores all sentences in the article according to their importance. The higher the score, the more important it is. It extracts several key sentences with the highest scores from the original text and forms a summary through sorting and reorganization.
[0003] In recent years, a lot of work has been done on the research of automatic text summarization technology at home and abroad. For example, the literature proposed an improved TextRank algorithm to optimize the weight of sentences in the text; Liu et al. applied BERT and Transformer to the extractive summarization task to realize the word vectorization of the input sequence, and achieved significant improvement in the Rouge index; Zhong et al. transformed the extractive summary into a semantic matching problem by comparing the summary methods at the sentence level and the summary level; Wang et al. proposed an extractive summary model based on heterogeneous graph networks, which achieved certain results in improving the readability of the summary; Sharma et al. proposed entity-driven The generative summary model uses entity information to generate a coherent summary containing useful information; Zhang et al. proposed a pre-trained generative summary model by extracting blank sentences, and used a new pre-training target for summary summarization and blank sentence generation; Zhu et al. proposed a fact-aware summary model, which extracts factual relationships from articles to construct knowledge graphs, and integrates them into the decoding process through neural graph computing; Chowdhury et al. proposed a hierarchical encoder based on structural attention to model inter-sentence and inter-document dependencies, which achieved significant improvements; Zheng et al. proposed a topic-aware abstract summary framework based on the potential semantic structure of documents represented by potential topics;
[0004] Generative summarization has high technical requirements for researchers, while extractive summarization, as a widely used automatic text summarization technology, has made some achievements. However, there is still room for improvement in extractive summarization in terms of processing the importance score of sentences in the text and redundant information.
[0005] In the extractive summary model, the whole sentence level extraction takes the whole sentence in the original text as the basic selection unit, and the reference summary is obtained by reorganizing the language based on the understanding of the original text; this whole sentence extraction usually contains both important and redundant information; in this regard, Zhou et al. studied the redundant information ratio in the model output summary; Xu et al. proposed a neural network model of post-extraction compression, which compresses the whole sentence after extracting the original text to remove redundant information in the whole sentence; Desai et al. proposed a method based on extraction compression, which modeled the compression summary model with authenticity and significance; Carbonell et al. proposed the maximum boundary correlation algorithm, which effectively avoided the problem of summary duplication; at the same time, some researchers proposed more fine-grained extraction units, such as words or phrases. Although these models can learn which words or phrases in the article are more important, it is still difficult to keep the semantic information intact when outputting the summary;
[0006] In addition, the training goal of the sentence-level extractive summarization model is to score independent sentences one by one to decide whether to use them as summaries. The model tends to extract sentences with the highest Rouge scores, but sentences with high Rouge scores are not necessarily in the best summaries. Through quantitative analysis, it is found that most of the best summaries are not composed of sentences with the highest sentence-level scores. In the CNN / Daily Mail dataset, only 18.9% of the best summaries come from sentences with high Rouge scores, which further shows that sentence-level extraction is prone to fall into local optimality and cannot obtain the best summary.
[0007] Therefore, there is an urgent need to design a clause-level automatic summarization model system based on deep learning, which can use clauses as extraction units and balance the importance and completeness of summary information at the same time to solve the problems existing in the above-mentioned prior arts. Summary of the invention
[0008] In view of the above-mentioned problems, the present invention aims to provide a clause-level automatic summary model system and a summary generation method based on deep learning. By taking clauses as extraction units, this model can balance the importance and completeness of summary information at the same time, effectively solve the redundancy problem in the sentence extraction process, and has the characteristics of high integrity of summary information and fast extraction speed.
[0009] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0010] A deep learning-based sentence-level automatic summarization model system, including a sentence-level extraction unit, a BERT-based encoder, a Transformer-based encoder, a classifier, and a BERTScore-based summary matcher;
[0011] The sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences;
[0012] The BERT-based encoder and the Transformer-based encoder are used to process the input sentence to obtain a vector representation C of the clause containing contextual semantic information. i ' , ' j ;
[0013] The classifier is used to classify the vector representation C processed by the BERT-based encoder and the Transformer-based encoder i ' , ' j Perform preliminary screening and output a summary of preliminary candidate clauses;
[0014] The BERTScore-based summary matcher is used to calculate the semantic similarity between the candidate clause summary and the original text, and obtain the summary that best matches the original text semantics.
[0015] Preferably, the sentence-level extraction unit is implemented using a heuristic algorithm based on dependency syntax;
[0016] The BERT-based encoder is implemented based on a BERT pre-trained model;
[0017] The Transformer-based encoder is implemented based on a Transformer model;
[0018] The classifier layer is a multi-layer perceptron;
[0019] The BERTScore-based summary matcher is implemented based on the automatic evaluation index BERTScore of deep learning.
[0020] A sentence-level automatic summary generation method based on deep learning, including the steps
[0021] S1. First, the sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences;
[0022] S2. Use the encoder based on the BERT pre-trained model and the Transformer model to process the input sentence constructed by the sentence-level extraction unit to obtain the vector representation C of the clause containing contextual semantic information i ' , ' j ;
[0023] S3. Use the classifier to preliminarily screen the vector representations processed by the BERT-based encoder and the Transformer-based encoder, and output preliminary candidate clause summaries;
[0024] S4. Finally, the semantic similarity between the candidate clause summary and the original text is calculated through the summary matcher based on BERTScore, and the summary that best matches the original text semantics is obtained, that is, the best summary.
[0025] Preferably, the calculation process of the sentence-level extraction unit described in step S1 includes:
[0026] (1) The whole sentence is parsed by a grammar analyzer to obtain the dependency relationship, which is expressed as a triple of "word A, label, word B";
[0027] Among them, the label represents the grammatical relationship between word A and word B;
[0028] (2) Split the entire sentence using the labels representing punctuation, conjunctions, and clause relations in the triples;
[0029] (3) merging units connected by special labels, including relative clause modifiers, adverbial clause modifiers, appositive modifiers, and clause complements;
[0030] (4) Determine whether the conjunction conj connects two clauses or two phrases. When the distance between the two connected elements is less than a fixed threshold, it is considered that the connection is two phrases and merged into one clause. Otherwise, it is considered that the connection is two clauses.
[0031] (5) A minimum unit length and a maximum unit length are predefined. When the unit length of an element is less than the minimum unit length, the element is merged with the previous element into a clause; otherwise, it is regarded as an independent clause.
[0032] Preferably, in the use of the BERT-based encoder described in step S2, BERT is used as the first layer encoder in the hierarchical encoder to read the input text and output the vector representation C′ of each clause in the original text. i,j .
[0033] Preferably, the BERT pre-trained model is used to output the vector representation C′ of each clause in the original text i,j The process includes
[0034] (1) For each clause C in the input document i,j , add the [CLS] tag at the beginning of the sentence to capture the clause features. The vector corresponding to this tag can be used for subsequent classification tasks, while for non-classification tasks, the [CLS] tag can be ignored; add the [SEP] tag at the end of the sentence to separate the clauses;
[0035] (2) Obtain the tag embedding, segment embedding, and position embedding of the given input respectively, and sum them up to form the final vector representation C′ i,j ;
[0036] Among them, tag embedding represents word vectors; segment embedding is used to distinguish between two sentences; position embedding represents the position information learned by the model.
[0037] Preferably, the calculation process of the Transformer-based encoder includes:
[0038] After passing through the BERT-based encoder, the clause vector representation C′ is obtained i,j Afterwards, in order to capture document-level features, a Transformer-based encoder is used for secondary encoding, and the representation is obtained through the Transformer's multi-head attention mechanism:
[0039]
[0040]
[0041] Among them, MHA(·) represents the multi-head attention mechanism in Transformer, LN(·) represents layer normalization, and FFN(·) represents a feedforward neural network containing two linear transformations.
[0042] Preferably, the calculation process of the classifier described in step S3 includes
[0043] After passing through the BERT and Transformer based encoder, a clause vector representation C″ containing document-level features is obtained. i,j , the input uses Sigmoid MLP classifier, which can map the output between (0,1) to indicate the probability of the predicted clause being extracted:
[0044] p(C″ i,j )=σ(W o C″ i,j +b o ) (5)
[0045] Where σ(·) represents the Sigmoid activation function, W o and b o represents a learnable parameter.
[0046] Preferably, the process of calculating the semantic similarity between the candidate clause summary and the original text by using the summary matcher based on BERTScore in step S4 includes:
[0047] (1) After the classifier outputs the candidate clause summary, the reference sentence x is matched with the candidate sentence. Calculate the recall rate, precision rate, and F1 value for each tag in the , and use the greedy algorithm to maximize the matching similarity score;
[0048] (2) At the same time, BERTScore introduces importance weighting, giving different weights to different words. Given M reference sentences The idf score of word w is:
[0049]
[0050] Where Γ(·) represents the indicator function;
[0051] (3) Update recall and precision using idf weights;
[0052]
[0053]
[0054] (4) Use BERTScore to calculate the semantic similarity between the candidate clause summary and the original text, and select the candidate clause summary with the highest summary level score as the final summary;
[0055] The summary score is calculated as:
[0056] score=F1 BERT (set(C i,j ),D) (12)
[0057] Among them, score represents the semantic score of the summary, set(C i,j ) represents the candidate clause summary, and D represents the original text.
[0058] Preferably, the recall rate, precision rate, and F1 value of the BERTScore described in step (1) are calculated as follows:
[0059]
[0060]
[0061]
[0062] Among them, x and Represents word x in the reference sentence and word x in the candidate sentence Contextual word embeddings from BERT.
[0063] The beneficial effects of the present invention are as follows: the present invention discloses a clause-level automatic summary model system and a summary generation method based on deep learning. Compared with the prior art, the improvements of the present invention are as follows:
[0064] (1) In view of the problems existing in the prior art, the present invention designs a clause-level automatic summarization model system and a summary generation method based on deep learning. The model includes a sentence-level extraction unit, an encoder based on BERT, an encoder based on Transformer, a classifier, and a summary matcher based on BERTScore. When in use, the sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences, and an encoder based on the BERT pre-trained model and the Transformer model is used to obtain a vector representation C′ of the clause containing contextual semantic information. i,j , and output preliminary candidate clause summaries through the classifier; finally, the semantic similarity between the candidate clause summaries and the original text is calculated through the summary matcher based on BERTScore, and the summary that best matches the original text semantics, i.e., the best summary, is obtained; experimental verification shows that the CS-ASum of the present invention is superior to the comparison model in both automatic evaluation and manual evaluation, and the Rouge index of the CS-ASum model is improved, indicating that the model of the present invention has the advantages of high integrity of summary information and fast extraction speed;
[0065] (2) To address the redundancy problem introduced by sentence-level extractive summarization and the semantic information integrity problem of word-level extractive summarization, the present invention proposes a clause-level extraction unit based on dependency syntax, and proposes a clause-level automatic summarization model system CS-ASum based on deep learning. This model effectively alleviates the redundancy problem while maintaining the integrity of the summary information.
[0066] (3) The present invention introduces a summary-level matcher based on BERTScore in CS-ASum, calculates the semantic matching score between the candidate summary and the original text, keeps the evaluation of the model in the training process consistent with the target evaluation, and obtains the summary that best matches the semantics of the original text. At the same time, experimental results show that the CS-ASum of the present invention can output better text summaries, and compared with the baseline model on the evaluation indicator Rouge, the output summary is concise, has low redundancy and is faithful to the original text, achieving better results. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is the structure of the sentence-level automatic summarization model system based on deep learning of the present invention.
[0068] Figure 2 This is an example of a dependency tree of the sentence-level extraction unit of the present invention. DETAILED DESCRIPTION
[0069] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0070] Embodiment 1: refer to the attached Figure 1-2 The deep learning-based clause-level automatic summarization model system (CS-ASum model) and summary generation method shown in the figure include a sentence-level extraction unit, a BERT-based encoder, a Transformer-based encoder, a classifier, and a BERTScore-based summary matcher;
[0071] The sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences. For example, the original input contains three whole sentences. The clause unit is constructed based on the dependency syntax. The first sentence and the third sentence are split into two clauses, namely C 1,1 , C 1,2 and C 3,1 , C 3,2 , the second sentence is an independent clause C 2,1 ;
[0072] The BERT-based encoder and the Transformer-based encoder are used to process the input sentence constructed by the sentence-level extraction unit to obtain a vector representation C′ of the clause containing contextual semantic information. i,j ;
[0073] The classifier is used to preliminarily screen the vector representations processed by the BERT-based encoder and the Transformer-based encoder, and output a preliminary candidate clause summary;
[0074] The BERTScore-based summary matcher is used to calculate the semantic similarity between the candidate clause summary and the original text, and obtain the summary that best matches the original text semantics, that is, the best summary.
[0075] Preferably, the sentence-level extraction unit uses a heuristic algorithm based on dependency syntax to split the whole sentence into clauses, and the specific calculation process includes:
[0076] (1) The whole sentence is parsed by the grammatical analyzer StanfordParser to obtain the dependency relationship, which is expressed as a triple of "word A, label, word B", where the label represents the grammatical relationship between word A and word B;
[0077] (2) Split the sentence using the labels representing punctuation (punct), conjunction (cc), and clause (mark) relations in the triples;
[0078] (3) In order to obtain more complete semantic units, some units connected by special tags are merged, including relative clause modifiers (acl:relcl), adverbial clause modifiers (advcl), appositive modifiers (appos) and clause complements (ccomp);
[0079] (4) Determine whether the conjunction conj connects two clauses or two phrases. When the distance between the two connected elements (the two clauses or phrases connected by the conjunction conj) is less than a fixed threshold, it is considered that the two phrases are connected and merged into one clause. Otherwise, it is considered that the two clauses are connected.
[0080] (5) Predefine the minimum unit length and the maximum unit length. When the unit length of an element is less than the minimum unit length, the element is merged with the previous element into a clause. Otherwise, it is regarded as an independent clause, that is, "It has created a virtual reality headset that works with any Android and iOS phone" and "is compatible with hundreds of virtual reality apps from the respective stores".
[0081] To more intuitively see how to split the sentence into clauses, Figure 2 An example dependency tree is shown; first, the sentence is split into three units; then, the first two units are merged because they are connected by the conjunction conj and do not exceed the threshold of being independent clauses; finally, the sentence is split into two clauses;
[0082] Table 1 shows the statistics of the CNN / Daily Mail dataset. On average, a sentence contains 1.63 clauses. Some short sentences in the dataset are considered as separate clauses, while some complex sentences can be split into 2 or 3 clauses.
[0083] Table 1: Average number and length of extraction units of different granularity
[0084] Extraction unit Average quantity Average length Sentence level 31.58 25.06 Clause Level 51.47 15.38
[0085] Preferably, the BERT-based encoder is implemented based on the BERT pre-training model. The essence of pre-training is a kind of transfer learning, which uses the parameters trained on a task as the shallow parameters for training a new task, and the high-level parameters are randomly initialized. When there is less training data, the pre-training method allows the model to learn based on a better initial state, improve the model training convergence speed, and achieve better performance. The BERT model uses the encoder structure in the Transformer to realize the extraction of bidirectional text information, thereby obtaining deeper feature information and mining deeper text features, and has achieved remarkable results in many natural language processing tasks.
[0086] BERT models generally have two structures: BERT-base and BERT-large. Different structures can be selected according to the actual usage scenarios and effects. When BERT is applied to specific downstream tasks, the model is usually fine-tuned. When facing specific problems, only the input and output need to be adjusted, and all parameters are fine-tuned end-to-end.
[0087] In the extractive text summarization task of this embodiment, the introduction of the BERT model in the embedding layer can provide powerful sentence embedding information for the entire model, which is also the key to improving the effect of extractive text summarization. To this end, this embodiment uses the BERT model to process the clause units constructed based on dependency syntax to realize the word vectorization of the input text. BERT is the first layer encoder in the hierarchical encoder, which reads the input text and outputs the vector representation C of each clause in the original text. i ' ,j , improve the encoding and feature extraction capabilities of the encoder;
[0088] The input representation of the BERT model is designed to explicitly represent a single sentence or sentence pair in a token sequence, while adding some special flags, such as [CLS], in order to successfully apply it to multiple downstream tasks; the steps to obtain the input sequence representation of the BERT model are:
[0089] (1) For each sequence (each clause C i,j ), put the special token [CLS] at the beginning of the sentence. The vector corresponding to this token can be used for subsequent classification tasks, while for non-classification tasks, the [CLS] token can be ignored; at the same time, put the special token [SEP] at the end of the sentence to separate sequences;
[0090] (2) Obtain the token embedding, segment embedding, and position embedding of the given input respectively, and sum them up to form the final input sequence;
[0091] Among them, tag embedding represents word vectors; segment embedding is used to distinguish between two sentences; position embedding represents the position information learned by the model;
[0092] The output representation of the BERT model includes two vector forms, character level and sentence level. The former uses vectors to represent all words in the input sequence; the latter uses the vector corresponding to the [CLS] flag in the model output to represent the semantic information of the entire sentence.
[0093] Specifically, for each clause C in the input document i,j , add [CLS] tag at the beginning of the sentence to capture the clause features, and add [SEP] tag at the end of the sentence to separate the clauses; given an article containing i clauses, use segment embedding to distinguish multiple clauses. For the i-th clause, assign segment embedding E according to whether i is odd or even. A or E B;exist Figure 1 In the vector C′ i,j The vector representing the corresponding [CLS] from the top BERT layer is used to extract the classification.
[0094] Preferably, the Transformer-based encoder is implemented based on the Transformer model, with the purpose of solving the problems of the recurrent neural network model that cannot be parallelized and cannot capture long-term dependencies well. The Transformer follows the encoder-decoder architecture, does not adopt an autoregressive model, and uses a self-attention mechanism and a fully connected layer to construct an encoder-decoder, so that the model can be trained in parallel while having global information.
[0095] The attention mechanism in Transformer is its core part. Its input consists of query, key and value pairs, and the output matrix is:
[0096]
[0097] The calculation process of the attention mechanism is actually to calculate the relationship between the source sequence and the target sequence. When the source sequence and the target sequence are the same, the attention calculation is the relationship within the sequence itself, that is, self-attention.
[0098] The multi-head attention in Transformer can further explore the internal relationship of the sequence, use h different linear spaces to project Q, K and V, and splice different attentions. The final multi-head attention value is:
[0099] MHA(Q,K,V)=Concat(H 1 ,H 2 ,...,H h )W O (2)
[0100] Among them, H i represents the attention matrix of each head, W O represents the mapping matrix;
[0101] Specifically in the CS-ASum model, after the BERT-based encoder obtains the clause vector representation C i ' ,j Afterwards, in order to capture document-level features, a Transformer-based encoder is used for secondary encoding, and the representation is obtained through the Transformer's multi-head attention mechanism:
[0102]
[0103]
[0104] Among them, MHA(·) represents the multi-head attention mechanism in Transformer, LN(·) represents layer normalization, and FFN(·) represents a feedforward neural network containing two linear transformations;
[0105] The specific process of the above C' calculation includes: C' is the output from BERT, C' is the input of the Transformer encoder, and after entering the Transformer encoder, the multi-head attention MHA(C') is first calculated in the first part, and then added to C' itself and layer normalization calculation is performed, that is, LN(C'+MHA(C')) to obtain the calculation result As input to the second part of the Transformer-based encoder, a feedforward neural network consisting of two linear transformations is first calculated Again with After adding them together, the layer normalization calculation is performed. Then we get the output C of the Transformer-based encoder, and the process is as follows Figure 2 shown.
[0106] Preferably, the classifier layer uses a multilayer perceptron (MLP) and uses the activation function Sigmoid, which is physically closest to biological neurons; after passing through the encoder based on BERT and Transformer, a clause vector representation C″ containing document-level features is obtained. i,j , the input uses Sigmoid MLP classifier, which can map the output between (0,1) to indicate the probability of the predicted clause being extracted:
[0107] p(C″ i,j )=σ(W o C″ i,j +b o )(5)
[0108] Where σ(·) represents the Sigmoid activation function, W o and b o represents a learnable parameter.
[0109] Preferably, after the classifier outputs the candidate clause summary, in order to obtain the best summary that best matches the original text semantics, it is necessary to calculate the semantic similarity between the candidate clause summary and the original text; in the natural language text generation task, the word overlap-based method calculates the word overlap between the reference sentence and the generated sentence, such as Rouge, which can only identify vocabulary changes and cannot perceive changes in the sentence semantic level, and has a large gap with manual evaluation; and the word vector-based method uses Word2vec or Sentence2vec to implement sentence vector representation, and uses cosine similarity to calculate relevance, but cannot fully capture the potential feature information in the sentence; therefore, the BERTScore-based automatic evaluation index based on deep learning is used to design a summary matcher based on BERTScore, and the BERT pre-trained model is used to extract the contextual features of the input words, realize the word vectorization of the input sentence, and obtain deeper text feature information; compared with general indicators, BERTScore has a higher correlation with manual evaluation and is better at extracting semantic information;
[0110] When calculating BERTScore, we match the reference sentence x with the candidate sentence Calculate the recall rate, precision rate, and F1 value for each tag in the BERTScore, and use the greedy algorithm to maximize the matching similarity score. The recall rate, precision rate, and F1 value of BERTScore are calculated as follows:
[0111]
[0112]
[0113]
[0114] At the same time, BERTScore introduces importance weighting, giving different weights to different words. Given M reference sentences The idf score of word w is:
[0115]
[0116] Among them, Γ(·) represents the indicator function. Since a single sentence is processed, the word frequency may be 1, so the complete tf-idf measurement is not applicable. The idf weight update formula (6) and (7) are used:
[0117]
[0118]
[0119] x and Represents the word x in the reference sentence and the word in the candidate sentence Contextual word embeddings from BERT;
[0120] In this embodiment, a summary matcher layer based on BERTScore is added after the classifier layer of the CS-ASum model. The five candidate clause summaries output by the classifier are reordered according to the original positions of the clauses in the original text. Considering that the average number of sentences in the reference summary is 4, the model output summary is set to contain 4 clauses. According to this number, the combination is performed to obtain Candidate clause summaries; calculated using BERTScore The semantic similarity between the candidate clause summaries and the original text is calculated, and the candidate clause summary with the highest summary level score is selected as the final summary;
[0121] The summary score is calculated as:
[0122] score=F1 BERT (set(C i,j ),D) (12)
[0123] Among them, score represents the semantic score of the summary, set(C i,j ) represents the candidate clause summary, and D represents the original text.
[0124] The implementation steps of the sentence-level automatic summarization model system based on deep learning in this embodiment include:
[0125] S1. First, the sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences;
[0126] S2. Use the encoder based on the BERT pre-trained model and the Transformer model to process the input sentence constructed by the sentence-level extraction unit to obtain the vector representation C of the clause containing contextual semantic information i ' ,j ;
[0127] S3. Use the classifier to preliminarily screen the vector representations processed by the BERT-based encoder and the Transformer-based encoder, and output preliminary candidate clause summaries;
[0128] S4. Finally, the semantic similarity between the candidate clause summary and the original text is calculated through the summary matcher based on BERTScore, and the summary that best matches the original text semantics is obtained, that is, the best summary.
[0129] Example 2: Different from the above-mentioned Example 1, the sentence-level automatic summarization model system based on deep learning described in Example 1 is verified using the following experiment and result analysis process:
[0130] 1. Dataset
[0131] The CNN / DailyMail dataset used in the experiment consists of news articles and their brief summaries. The statistical information before (whole sentence level) and after (sub-sentence level) data preprocessing is shown in Table 2.
[0132] Table 2: Statistics of the CNN / DailyMail dataset before (sentence level) and after (clause level) processing
[0133] CNN / DailyMail Training set Validation set Test Set Number of documents 287227 13368 11490 Average number of complete source text sentences 31.58 26.72 27.05 Average source text length 791.52 769.44 778.45 Average number of complete sentences in reference abstracts 3.79 4.11 3.88 Average reference abstract length 55.18 61.45 58.33 Average number of source text clauses 51.47 50.50 50.47
[0134] 2. Hyperparameter Settings
[0135] The experiment uses the Adam optimizer, the learning rate is set to α = 2e-5, and the momentum parameter is set to β 1 =0.9,β 2 =0.999, the residual is set to ε=10 -8 ; The model is implemented using PyTorch and the bert-base-uncased version of BERT, which contains 12 Transformer layers; the model is trained for 50,000 steps with a batch-size of 32. After training for 10,000 steps, the model is saved and evaluated every 1,000 steps; the three best checkpoints on the validation set are used to record the best model; Dropout is set to 0.1;
[0136] 3. Evaluation Method
[0137] The automatic evaluation method uses the general summary evaluation criterion Rouge. The Rouge indicator measures the summary quality by calculating the n-gram phrase overlap between the reference summary and the candidate summary, and is mainly composed of Rouge-N (N can be 1, 2, 3, etc.) and Rouge-L. This embodiment uses Rouge-1 (unigram) and Rouge-2 (bigram) to measure the richness of summary information. Rouge-L represents the longest common subsequence between the reference summary and the candidate summary, and is used to measure the fluency of the summary content.
[0138] At the same time, in order to deeply analyze the similarity of the abstracts at the semantic level and further evaluate the quality of the model output summary, the experiment also conducted manual evaluation;
[0139] 4. Comparison Models
[0140] In order to verify the effectiveness of the model in Example 1 of the present invention, the comparison model in the experiment is as follows:
[0141] (1) Lead-3: extract the first three sentences of the original text as the summary;
[0142] (2) PoinGenCov: Based on the standard sequence-to-sequence architecture, it introduces a coverage mechanism to improve problems such as inaccurate information generation;
[0143] (3) FastAbsRL: A two-stage summarization model based on a sentence-level policy gradient algorithm that combines the advantages of extractive and generative methods;
[0144] (4) JECS: A single-document neural summarization model that jointly extracts and compresses data, treating the summarization problem as a series of local decisions;
[0145] (5)TwoStageRL: A two-stage summarization model based on an encoder-decoder architecture;
[0146] (6) BERTAbsRL: An extraction system that integrates the BERT model and globally optimizes the summary-level Rouge score through reinforcement learning;
[0147] (7) BERTExtAbs: uses a standard encoder-decoder framework, combined with pre-trained and Transformer decoders, and fine-tuned on extractive and generative summarization tasks, respectively;
[0148] (8) UniLM: Unified pre-trained language model that completes generative summarization tasks through fine-tuning on downstream tasks;
[0149] (9) PEGASUS: A pre-trained model that uses extracted interstitial sentences for summarization. It can be pre-trained for generation tasks and the model can autonomously learn high-importance sentences in the original text.
[0150] (10) TextRank: A graph-based ranking model where each sentence corresponds to a node. The weights of adjacent edges represent the semantic similarity of the nodes. Multiple nodes with the highest weights are extracted to form the final summary.
[0151] (11) NeuSum: An end-to-end neural network framework for extractive text summarization by jointly learning sentence scoring and sentence selection.
[0152] (12) HiBERT: The sentence-level encoder learns sentence representations through intra-sentence information, and the paragraph-level encoder learns sentence representations with contextual information through inter-sentence information;
[0153] (13) HeterGraph: An extractive summarization model based on heterogeneous graphs. It introduces different types of nodes into graph-based neural networks, including word nodes and sentence nodes. The model has good scalability.
[0154] (14) DiscoBERT: A discourse-aware extractive summarization model that captures long-term dependencies through graph convolutional neural network encoding.
[0155] (15) CUPS: proposed a compression summarization system that decomposes cross-level compression into two learnable objectives: rationality and saliency.
[0156] (16) MatchSum: Converts the extractive summarization task into a semantic matching problem and directly performs summary-level extraction.
[0157] (17) MatchSum+CUPSCMP: a compression module that combines the MatchSum model and the CUPS model;
[0158] (18) A2C-RLAS: A reinforced automatic summarization model system based on the dominant actor-critic algorithm, connecting the extractor and rewriter through reinforcement learning, and optimizing the model based on the semantic reward of BERTScore;
[0159] (19) CS-ASum: the model in this paper;
[0160] 5. Automatic evaluation and analysis
[0161] The comparison results of Rouge indicators of different models are shown in Table 3. In the table, R-1, R-2, and RL scores represent Rouge-1, Rouge-2, and Rouge-L, respectively. The R-AVG score refers to the average of Rouge-1, Rouge-2, and Rouge-L. The highest score of the indicator is marked in bold;
[0162] Table 3: Comparison of Rouge evaluation indicators of different models
[0163]
[0164] As can be seen from Table 3, the proposed model CS-ASum achieved the highest scores in Rouge-1, Rouge-2 and Rouge-L metrics, indicating that CS-ASum surpassed all the comparison models, and the model performed best overall, and obtained a higher quality summary, proving the effectiveness of the proposed model CS-ASum.
[0165] Compared with the generative summary model PoinGenCov, CS-ASum improved by 5.53, 4.89 and 5.42 percentage points in Rouge-1, Rouge-2 and Rouge-L respectively. Such a performance gap may be because PoinGenCov uses a single-layer Bi-RNN as an encoder, which works well for short texts, but when processing long texts, the encoder cannot correctly capture the semantic features of the text and has a long-distance dependency problem. At the same time, the decoder based on a single-layer neural network has insufficient semantic representation memory, resulting in poor quality of summaries generated by the model.
[0166] Compared with the extractive summarization model NeuSum, CS-ASum improved by 3.47, 3.16, and 3.82 percentage points in Rouge-1, Rouge-2, and Rouge-L indicators, respectively. This may be because although the NeuSum model considers the sentence features of the source text and the influence between the context sentences, it does not fully consider the potential semantic relationship between words and the semantic features of the sentences in the text, resulting in no improvement in model performance.
[0167] Compared with the A2C-RLAS model, CS-ASum achieved 3.73, 3.96 and 2.91 percentage point improvements in Rouge-1, Rouge-2 and Rouge-L indicators respectively; the R-AVG score of the CS-ASum model is 36.34, which is 3.53 percentage points higher than that of the A2C-RLAS model. This shows that it is very beneficial to use the pre-trained model for extractive summarization tasks, as it provides powerful sentence embedding information and context information, making the extracted sentences more accurate.
[0168] 6. Manual evaluation analysis
[0169] To ensure the robustness of the model in this paper, the experiment also conducted manual evaluation. Referring to the work of Bae et al., the performance of the summary system was evaluated by focusing on the relevance and readability of the model output summary. Relevance means that the output summary contains important information in the original text and avoids irrelevant or repeated redundant information; readability means that the output summary is fluent, grammatically correct and coherent;
[0170] 100 articles were randomly sampled from the CNN / Daily Mail test set. Considering that manual evaluation requires a lot of time, three human testers were asked to sort the output summaries of the two models, BERTAbsRL and CS-ASum, according to relevance and readability, and score the summaries 2, 1, and 0 respectively according to the rankings. The models were anonymized and randomly shuffled. Before this work, in addition to the three model output summaries, the original and reference summaries were also shown to the human testers. The scoring results of the model output summaries are shown in Table 4, where the highest scores are marked in bold.
[0171] Table 4: Manual evaluation results
[0172] Model Relevance readability Total score BERTAbsRL 58 60 118 CS-ASum 72 63 135
[0173] 7. Ablation Experiment
[0174] In order to experimentally verify the necessity of each layer structure in the CS-ASum model, this embodiment performs ablation analysis on the clause extraction, BERT pre-training and BERTScore summary matcher in the CS-ASum model. The results are shown in Table 5. The highest score is marked in bold.
[0175] Table 5: Ablation experiment results
[0176]
[0177]
[0178] As can be seen from Table 5, compared with CS-ASum based on sentence-level extraction, CS-ASum based on clause-level extraction improves by 5.39, 3.74, and 5.88 percentage points in Rouge-1, Rouge-2, and Rouge-L indicators, respectively, indicating that clause-based construction can provide more accurate training labels for model training;
[0179] To verify the impact of pre-training on the performance of the summary model, this paper conducts ablation analysis on the segment embedding and position embedding in CS-ASum. After removing the segment embedding, the model's Rouge-1, Rouge-2, and Rouge-L indicators dropped by 0.21, 0.21, and 0.25 percentage points, respectively, indicating that it is challenging for the model to understand the hierarchical structure when there is no difference in information from different sources. After removing the position embedding, the performance of the model dropped significantly, dropping by 3.72, 3.04, and 3.75 percentage points, respectively, on the Rouge-1, Rouge-2, and Rouge-L indicators, indicating that position information has a greater impact on the extractive summary model.
[0180] Compared with the model without summary matcher, CS-ASum with summary matcher improves Rouge-1, Rouge-2 and Rouge-L indicators by 1.95, 2.00 and 1.79 percentage points respectively, indicating that the model can output the summary with the highest summary-level semantic score by using summary matcher to semantically match candidate clause summaries with the original text;
[0181] 8. Summary Example Analysis
[0182] Table 6 shows the summary samples of the proposed model CS-ASum and other comparison models. The precision (P), recall (R) and F1 value (F1) of CS-ASum in BERTScore are 99.2779, 99.5608 and 99.4191 respectively, which achieves the best results compared with other comparison models. This shows that the semantics of the summary output by CS-ASum is almost consistent with that of the reference summary, achieving the purpose of text summarization.
[0183] Table 6: Summary examples of different models
[0184]
[0185]
[0186] As can be seen from Table 6, the output summary content of the extractive summary model Lead-3 cannot correctly match the reference summary content. It can be seen that the method based on simple rules has obvious defects when the topic of the article is not given in the first three sentences of the original text; MatchSum successfully outputs three summaries that match the reference summary, but ignores an important sentence of the original text, which means that the model does not fully understand the semantic information of the original text, resulting in the loss of key information in the output summary;
[0187] The first sentence in the summary output by the generative summary model PoinGenCov matches the first sentence in the reference summary, indicating that the model only focuses on one important sentence in the article and ignores other important sentences. FastAbsRL generates a five-sentence summary, but only the second sentence matches the important sentence in the original text. The generated summary contains too much redundancy and is inconsistent with the original text.
[0188] The model CS-ASum in this paper outputs a four-sentence summary, which focuses on the key sentences in the same position in the article as the reference summary. It is more comprehensive and highly close to the reference summary, indicating the superiority of CS-ASum.
[0189] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A sentence-level automatic summarization model system based on deep learning, Features: Includes sentence-level extraction unit, BERT-based encoder, Transformer-based encoder, classifier, and BERTScore-based summary matcher; The sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences; The BERT-based encoder and the Transformer-based encoder are used to process the input sentence to obtain a vector representation C of the clause containing contextual semantic information. i ' , ' j ; The classifier is used to classify the vector representation C processed by the BERT-based encoder and the Transformer-based encoder i ' , ' j Perform preliminary screening and output a summary of preliminary candidate clauses; The BERTScore-based summary matcher is used to calculate the semantic similarity between the candidate clause summary and the original text, and obtain the summary that best matches the original text semantics; wherein: The sentence-level extraction unit is implemented using a heuristic algorithm based on dependency syntax; The BERT-based encoder is implemented based on a BERT pre-trained model; The Transformer-based encoder is implemented based on a Transformer model; The classifier layer is a multi-layer perceptron; The BERTScore-based summary matcher is implemented based on the automatic evaluation index BERTScore of deep learning.
2. A sentence-level automatic summary generation method based on deep learning, implemented using the automatic summary model system as claimed in claim 1, Features: Included steps S1. First, the sentence-level extraction unit is used to split the original data into words and sentences to construct input words and sentences; S2. Use the encoder based on the BERT pre-trained model and the Transformer model to process the input sentence constructed by the sentence-level extraction unit to obtain the vector representation C of the clause containing contextual semantic information i ' , ' j ; S3. Use the classifier to preliminarily screen the vector representations processed by the BERT-based encoder and the Transformer-based encoder, and output preliminary candidate clause summaries; S4. Finally, the semantic similarity between the candidate clause summary and the original text is calculated through the summary matcher based on BERTScore, and the summary that best matches the original text semantics is obtained, that is, the best summary; The calculation process of the sentence-level extraction unit described in step S1 includes: (1) The whole sentence is parsed by a grammar analyzer to obtain the dependency relationship, which is expressed as a triple of "word A, label, word B"; Among them, the label represents the grammatical relationship between word A and word B; (2) Split the entire sentence using the labels representing punctuation, conjunctions, and clause relations in the triples; (3) merging units connected by special labels, including relative clause modifiers, adverbial clause modifiers, appositive modifiers, and clause complements; (4) Determine whether the conjunction conj connects two clauses or two phrases. When the distance between the two connected elements is less than a fixed threshold, it is considered that the connection is two phrases and merged into one clause. Otherwise, it is considered that the connection is two clauses. (5) predefine a minimum unit length and a maximum unit length. When the unit length of an element is less than the minimum unit length, the element is merged with the previous element into a clause, otherwise it is regarded as an independent clause; The process of calculating the semantic similarity between the candidate clause summary and the original text by using the summary matcher based on BERTScore in step S4 includes: (1) After the classifier outputs the candidate clause summary, the reference sentence x is matched with the candidate sentence. Calculate the recall rate, precision rate, and F1 value for each tag in the , and use the greedy algorithm to maximize the matching similarity score; (2) At the same time, BERTScore introduces importance weighting, giving different weights to different words. Given M reference sentences The idf score of word w is: Where Γ(·) represents the indicator function; (3) Update recall and precision using idf weights; (4) Use BERTScore to calculate the semantic similarity between the candidate clause summary and the original text, and select the candidate clause summary with the highest summary level score as the final summary; The summary score is calculated as: score=F1 BERT (set(C i,j ),D)(12) Among them, score represents the semantic score of the summary, set(C i,j ) represents the candidate clause summary, and D represents the original text.
3. The method for generating a clause-level automatic summary based on deep learning according to claim 2, Features: In the use of the BERT-based encoder described in step S2, BERT is used as the first layer encoder in the hierarchical encoder to read the input text and output the vector representation C of each clause in the original text. i ' ,j .
4. The sentence-level automatic summary generation method based on deep learning according to claim 3, Features: Use the BERT pre-trained model to output the vector representation C of each clause in the original text i ' ,j The process includes (1) For each clause C in the input document i,j , a [CLS] tag is added to the beginning of the sentence to capture the clause features. The vector corresponding to this tag is used for subsequent classification tasks, while for non-classification tasks, the [CLS] tag is ignored; Add [SEP] tag at the end of the sentence to separate clauses; (2) Obtain the tag embedding, segment embedding, and position embedding of the given input respectively, and sum them up to form the final vector representation C i ' ,j ; Among them, tag embedding represents word vectors; segment embedding is used to distinguish between two sentences; position embedding represents the position information learned by the model.
5. The sentence-level automatic summary generation method based on deep learning according to claim 4, Features: The calculation process of the Transformer-based encoder includes: After passing through the BERT-based encoder, the clause vector representation C is obtained. i ' ,j Afterwards, in order to capture document-level features, a Transformer-based encoder is used for secondary encoding, and the representation is obtained through the Transformer's multi-head attention mechanism: Among them, MHA(·) represents the multi-head attention mechanism in Transformer, LN(·) represents layer normalization, and FFN(·) represents a feedforward neural network containing two linear transformations.
6. The method for generating a clause-level automatic summary based on deep learning according to claim 5, Features: The calculation process of the classifier described in step S3 includes: After passing through the BERT and Transformer based encoder, a clause vector representation C containing document-level features is obtained. i ' , ' j , the input uses Sigmoid's MLP classifier, and the output is mapped between (0,1) to indicate the probability of the predicted clause being extracted: p(C i ' , ' j )=σ(W o C i ' , ' j +b o ) (5) Where σ(·) represents the Sigmoid activation function, W o and b o represents a learnable parameter.
7. The method for generating a clause-level automatic summary based on deep learning according to claim 6, Features: The recall rate, precision rate, and F1 value of the BERTScore described in step (1) are calculated as follows: Among them, x and Represents the word x in the reference sentence and the word in the candidate sentence Contextual word embeddings from BERT.
Citation Information
Patent Citations
Two-stage text abstraction method
CN112100365A