A long text extractive summary generation method and device based on hierarchical iteration

Through the method based on hierarchical iteration, the extraction abstract is generated for long text, which solves the problems of large computing resources, semantic loss and lack of structural modeling of long text extraction abstracts, and achieves better text understanding and analysis capabilities.

CN118332101BActive Publication Date: 2025-05-16INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410400400.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-03
Publication Date
2025-05-16
Estimated Expiration
2044-04-03

AI Technical Summary

Technical Problem

When faced with long text, the existing extracted abstract technology consumes a lot of computing resources, and there are problems such as semantic loss and lack of long text structure modeling.

Method used

A long text extraction abstract generation method based on hierarchical iteration is adopted. By obtaining word vectors, position vectors and structure subtitle vectors, semantic encoding is performed, and semantic information is transferred layer by layer from sentence level to document level along the text structure route. Then, iterative update is performed again from document level to sentence level, and the hidden layer representation of each level is obtained. Finally, the optimal summary sentence is selected by fusing the hidden layer representation of each level.

Benefits of technology

It effectively overcomes the problems of large computing resources, semantic loss and lack of structural modeling of long text extraction abstracts, and improves the model's understanding and analysis ability of long texts, which is better than previous baseline models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118332101B_ABST
    Figure CN118332101B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of text information extraction, and relates to a method and device for generating extractive summaries of long texts based on hierarchical iteration. The method comprises: obtaining word vectors, position vectors, and structural subtitle vectors of characters in the text, adding them up as input for semantic coding, using a long text pre-trained language model as a semantic encoder, and performing semantic coding; sending the vectors after semantic coding to encoders at each level, transferring the semantic information hierarchically from the sentence level to the document level along the text structure route, and then transferring the semantic information hierarchically from the document level to the sentence level again, realizing iterative updating, and obtaining the hidden layer representations of each level; comprehensively evaluating each sentence by fusing the hidden layer representations of each level, and selecting the optimal summary sentence. The present invention can overcome the problems of large consumption of computing resources, semantic loss, and lack of long text structure modeling in the existing extractive summaries for long texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of text information extraction, and relates to several technical methods for extracting text summaries from article-level data, specifically a long text extraction summary generation method and device based on hierarchical iteration. Background Art

[0002] At present, automatic summarization technology refers to the use of computers to automatically extract simple and coherent short texts from original documents that can fully and accurately reflect the core content of the document. It is more efficient than manual summarization. With the help of automatic summarization technology, we can deeply understand and analyze massive web data, automatically generate summaries of long texts, reduce the complexity of the content, reduce the structural differences between texts, and greatly reduce labor costs. This has high research value and application needs in information retrieval, content filtering, public opinion analysis, situation awareness and other fields. Research on extractive summarization technology for long texts helps to improve the model's understanding and analysis of long texts, which is one of the difficulties in natural language processing and has high research value in natural language understanding tasks.

[0003] Text summarization technology can be mainly divided into generative summarization and extractive summarization. At present, the extractive summarization technology is more mature, and the generated sentences are more fluent and reliable. Therefore, this technology intends to adopt the extractive summarization technology route. The current mainstream methods are all based on deep learning, including extractive summarization based on sequence models, extractive summarization based on graph neural networks, extractive summarization based on reinforcement learning, and extractive summarization based on semantic matching.

[0004] 1. Extractive Summarization Based on Sequence Model

[0005] Sequence models are the most commonly used model structures for modeling natural language texts, including CNN, LSTM, Transformer, etc. In extractive summarization, SummaRuNNer first modeled it as a sequence labeling task. It first converts the word vectors obtained through Word2Vec training into hidden layer representations, and then constructs a two-layer bidirectional GRU network. It first updates the word vectors in the sentence, aggregates the sentence vectors, and then inputs the sentence vectors into the sentence-level GRU network to interact with the semantic relationship between sentences. Finally, when predicting the summary sentence at the classification layer, it considers multiple aspects of information such as importance and novelty to extract the summary sentence. This work is an early work on extractive summarization based on deep learning, which has a great influence and inspiration on subsequent work.

[0006] With the popularity of Transformer models and the advent of large-scale pre-trained models, the effect of sentence semantic encoding has been significantly improved. Among them, BertSum first introduced pre-trained models into the extractive summary task. It inserts [CLS] and [SEP] identifiers at the beginning and end of each sentence in the text, respectively, and uses the vector of [CLS] as the vector representation of the sentence. Then, the sentence vector sequence is input into the two-layer Transformer network to model the relationship between sentences. Finally, the summary sentence is extracted based on the hidden vector of the sentence, achieving excellent results. At the same time, the author proposed the Trigram-Blocking algorithm in the paper to post-process the generated summary sentence: if there is a 3-gram in common between the to-be-selected sentence and the selected sentence, the sentence will not be selected. This algorithm can avoid the redundancy between summary sentences to a certain extent, especially on the CNN / DM dataset, and has achieved good results.

[0007] This type of extractive summary model based on the sequence model is intuitive and efficient, but it is difficult to capture complex inter-sentence relationships based on the sequence model.

[0008] 2. Extractive Summarization Based on Graph Neural Network

[0009] The graph-based ranking method is a classic model in extractive summarization, among which the TextRank algorithm based on PageRank has been widely used. Although it has been surpassed by the deep learning-based method in terms of effect, its idea of ​​using graphs to capture complex inter-sentence relationships can still inspire subsequent work. With the popularity of graph neural networks, some models have also begun to use graph neural networks to model inter-sentence relationships. HeterGraphSum uses words and sentences in a document as nodes of the graph, where word nodes are initialized using word vectors and word encoders, and sentences are initialized using word vectors and CNN and BiLSTM encoders. Sentences are connected to the word nodes they contain, and the edge information is initialized using Tf-idf values. Then GAT is used to iteratively update sentence nodes and word nodes (first use sentence nodes to aggregate the information of neighboring word nodes, and then word nodes aggregate the information of neighboring sentence nodes), and finally select summary sentences based on sentence nodes. It has achieved excellent results without using a pre-trained model.

[0010] HAHSum is also based on graph neural networks, but it focuses more on modeling redundant information between sentences. It uses ALBERT to semantically encode the original document, and then uses an abstract layer to learn word-level information interactions between different sentences, including word nodes, sentence nodes, and entity nodes. After the abstract layer aggregates information from words and entities to sentence nodes, it uses a redundant layer to learn redundant information between different sentences, connects sentences with triple-repetition relationships, and more accurately models redundancy between sentences through information transfer in heterogeneous graphs. Node updates in heterogeneous graphs are all completed using GAT. For modeling redundancy, the author designed an iterative update method and used a gating mechanism to avoid smoothing problems in GNNs.

[0011] The graph neural network-based solution can capture relatively complex inter-sentence relationships and has more advantages than the sequence model when modeling inter-sentence relationships. However, its structure needs to be designed manually, and the introduced prior knowledge may not be correct, which may have some impact on the performance of the model.

[0012] Feature-based malicious domain name detection methods require corresponding domain knowledge because they need to extract feature information, and the designed features are easily modified by attackers.

[0013] 3. Extractive Summarization Based on Reinforcement Learning

[0014] Since extractive summarization is modeled as a sequence labeling task, the process is supervised by the 0 or 1 labels that are previously assigned to sentences by the greedy algorithm. During model training, the general objective function optimized is the cross entropy value between the sentence score and the 0 or 1 label. However, when verifying the model effect, the evaluation indicator is the ROUGE value between the generated summary and the manual summary, which causes the inconsistency between the optimization goal and the evaluation indicator. To address this problem, some researchers introduced reinforcement learning to try to reduce the inconsistency between the optimization goal and the evaluation indicator.

[0015] RankSum conceptualizes the extractive summarization problem as a sentence ranking task, and globally optimizes the ROUGE evaluation index through the learning goal of reinforcement learning. It first uses a convolutional neural network as a sentence encoder, aggregates word vectors into sentence vectors through convolution calculations and maximum pooling, then uses a document encoder with an RNN structure to extract features, and finally extracts sentences as summaries based on these features. During the training process, it does not use cross entropy loss combined with binary labels as training objectives, but introduces reinforcement learning technology, uses the ROUGE value between sentences and summaries as rewards, and combines the probability of a sentence becoming a summary as a reinforcement learning training strategy for gradient descent learning.

[0016] BanditSum regards extractive summarization as a contextual bandit problem, extracts summary candidates from the original document by random sampling, and uses the ROUGE value of the candidate and manual summaries as the reward function to train the model.

[0017] Although reinforcement learning makes the training objectives closer to the evaluation indicators, its training process is more difficult and the effect is not very significant. This issue still deserves further exploration.

[0018] 4. Extractive Summarization Based on Semantic Matching

[0019] Recently, some work has not been confined to the previously proposed sequence tagging framework, but has migrated the paradigm of extractive summarization to the text matching task. MatchSum encodes the candidate summary sentence set and the original text through a pre-trained model to obtain a semantic vector, calculates the cosine similarity between the two, and selects the most similar candidate summary set. During the training process, it not only shortens the distance between the candidate summary and the original text and the artificial summary, but also arranges the candidate summary in descending order according to the ROUGE value between them and the artificial summary. A RankLoss is designed to increase the distance between different candidate summaries and enhance their semantic representation, and finally achieves excellent results. One of the reasons for the success of this framework is that compared with the sequence tagging model, it can comprehensively select summary sentences from the summary level instead of considering it from the sentence level. However, for an original document to be generated, the number of sentences is often relatively large. Therefore, during the training process, the number of candidate summary sets generated directly is too large, and there is a dimensionality disaster. The paper uses the BertSum model to screen sentences in advance to solve this problem. However, this makes the model a two-stage training task, which cannot be trained end-to-end, increases the difficulty of training and deployment, and has the problem of error accumulation. Summary of the invention

[0020] In order to overcome the problems of large consumption of computing resources, semantic loss and lack of long text structure modeling in existing extractive summarization for long texts, the present invention proposes a long text extractive automatic summary generation method and device based on hierarchical iteration.

[0021] The technical solution adopted by the present invention is as follows:

[0022] A long text extractive summary generation method based on hierarchical iteration includes the following steps:

[0023] Get the word vector, position vector and structure subtitle vector of the characters in the text;

[0024] The word vector, position vector and structure subtitle vector are added together as the input of semantic encoding, and the long text pre-trained language model is used as the semantic encoder for semantic encoding;

[0025] The semantically encoded vectors are sent to encoders at each level, and the semantic information is transferred hierarchically from the sentence level to the document level along the text structure route, and then transferred hierarchically from the document level to the sentence level again, to achieve iterative updates and obtain the hidden layer representations of each level;

[0026] By integrating the hidden representations of each level, each sentence is comprehensively evaluated and the optimal summary sentence is selected.

[0027] Furthermore, the position vector includes a position vector representing a linear sequence and a position vector including hierarchical structure information.

[0028] Furthermore, the long text pre-trained language model is a Longformer model, and the attention mechanism of the Longformer model includes local attention and global attention; each word in the sequence performs calculations of the two attention mechanisms. First, each word performs a complete self-attention calculation on words of a preset window size, which is local attention. Then, each word in the sequence performs attention calculation on special characters, which is global attention. Finally, the results of the two attention mechanisms are combined to obtain the final representation of the word.

[0029] Furthermore, the Longformer model adopts the following steps to obtain the initialization representation of the sentence: inserting [CLS] characters at the beginning of each sentence, inserting [SEP] characters at the end of the sentence, and then setting [CLS] as a special character for global attention in Longformer, then inputting all characters into Longformer to obtain its hidden layer representation, and introducing a weighted attention averaging mechanism to obtain the hidden layer representation of the sentence node.

[0030] Furthermore, the encoders at each level are constructed by BiLSTM, and for the jth paragraph in the i-th layer The hidden layer representation of the text is constructed, and Bi-LSTM is used as a hierarchical encoder to capture the dependencies between them. The hidden layer representation is then mean-pooled and a feed-forward neural network is used to obtain the corresponding hidden layer representation of the jth paragraph of the i-th layer. The hidden layer representations from low to high levels are all modeled as a sequence, and the information is transmitted hierarchically from bottom to top according to the encoders of each layer until the vector representation of the document is obtained.

[0031] Furthermore, the layered transfer from the document level to the sentence level is performed again to achieve iterative updating, including: introducing an attention mechanism, for the i-th level encoder of the k-th iteration, using the hidden layer vector representation corresponding to the i-th layer in the k-1th round as the query value, and the hidden layer vector representation corresponding to the i-1th layer obtained in the k-1th round as the key value, and calculating the initial state vector of the BiLSTM.

[0032] Furthermore, during the model training process, an auxiliary training goal of paragraph attribution is designed for each sentence: the substructure titles with similar meanings in the text are grouped into one category, so as to predict which category each sentence belongs to; then the summary sentence selection task and the sentence category prediction task are combined as the training goal of the entire model.

[0033] A long text extractive summary generation device based on hierarchical iteration, comprising:

[0034] The preprocessing module is used to obtain the word vector, position vector and structure subtitle vector of the characters in the text;

[0035] The text semantic encoding module is used to add the word vector, position vector and structure subtitle vector as the input of semantic encoding, and use the long text pre-trained language model as the semantic encoder to perform semantic encoding;

[0036] The information transfer module is used to send the semantically encoded vectors to the encoders at each level, transfer the semantic information hierarchically from the sentence level to the document level along the text structure route, and then transfer it hierarchically from the document level to the sentence level again, to achieve iterative update and obtain the hidden layer representation of each level;

[0037] The information fusion selection module is used to comprehensively evaluate each sentence by fusing the hidden layer representations of each level and select the optimal summary sentence.

[0038] The beneficial effects of the present invention are as follows:

[0039] The present invention proposes a long text extractive summary generation scheme based on hierarchical iteration, which hierarchically models the complex text structure of the long text to complete the hierarchical transmission of information, and then models the text multiple times by iteratively updating the representation vectors of each level to improve the model understanding effect. The results on the existing data set show that the model of the present invention is better than the previous baseline model, and verify the rationality of the model design of the present invention and the effectiveness of the parameter setting. The present invention can overcome the problems of large computing resource consumption, semantic loss and lack of long text structure modeling when the existing extractive summary is for long text. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the overall architecture of the present invention.

[0041] Figure 2 is a schematic diagram of information transmission details, where K th Represents the input of the kth iteration.

[0042] Figure 3 It is the impact of the number of iterations on the model effect. DETAILED DESCRIPTION

[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below through specific embodiments and drawings.

[0044] The main contents of the present invention include:

[0045] 1) The present invention introduces hierarchical structure information in long texts. Before the text sequence is input for semantic encoding, in addition to generating a position vector representing a linear sequence, a position vector containing hierarchical structure information is generated to provide position information for the current character.

[0046] 2) Longformer, a long text pre-trained language model, is introduced as a text semantic encoder. Local sparse attention is set to each character (token) and global attention is set to CLS characters to achieve a balance between computing resource consumption and semantic encoding effect.

[0047] 3) The semantically encoded vector is sent to the encoders at each level built by BiLSTM, and the semantic information is transmitted layer by layer from bottom to top along the text structure.

[0048] 4) This transfer process is used as an iterative unit to perform attention calculation on the semantic representation vector of the previous round and the semantic vector representation of the next level. The result is used as the initialization vector of the BiLSTM layer to achieve the goal of iterative update.

[0049] 5) Finally, the hidden layer representations of each level are comprehensively considered to select the optimal summary sentence.

[0050] The present invention first models the extractive summarization problem as a classic sequence labeling problem. Given a document D containing N sentences, D = {s1, s2, ..., s N}, here s i represents the i-th sentence in document D, where s i Contains M words {w i1 ,w i2 ,...,w iM}. For sentence s i First, use the greedy strategy-based Oracle algorithm to label it with a binary label y i∈{0,1}, that is, the sentences are divided into positive and negative samples, where the positive samples represent summary sentences and the negative samples represent non-summary sentences. The extractive summarization model predicts a score for each sentence as the probability of the summary sentence, and finally determines the final summary sentence set based on the probability of each sentence and the number of summary sentences required. The HISum (Hierarchical Modeling and Iterative Representation for Extractive Summarization) model proposed in this invention will solve this problem definition. The overall framework of the present invention is as follows Figure 1 shown.

[0051] 1) Hierarchical modeling framework

[0052] First, the hierarchical structure of the text is defined. The minimum hierarchical structure of an article is set to sentence S, and the maximum hierarchical structure is set to article D. There are several layers of structure (sentences, sections, chapters, etc.) in between. L = {L 1 ,L 2 ,...,L n}, for the text paragraph at level i have in represents the mth sub-paragraph in the jth paragraph. In extractive summarization, the selection of summaries is performed at the smallest level, i.e., the sentence level. Before inputting the word sequence into the pre-trained language model for semantic encoding, its position information needs to be encoded first. In addition to calculating the position vector PE as a sequence of characters word , and also take the structure of each level as a sequence and calculate the position vector PE L , for those in The characters in have the same i-th level position vector The final position vector PE is obtained by summing the position vectors of the above layers. Here, the position encoding calculation method used in Transformer is used to calculate and encode the positions of the above two layers:

[0053]

[0054]

[0055] Here pos represents the position, i represents the dimension in the vector, and d model Represents the dimension of the hidden layer representation.

[0056] When dividing the text into hierarchical structures, a title is usually set for a part of the content, which often represents the core semantics of the content. Here, it is directly converted into the corresponding word vector, called the structural subtitle vector LE. Then the word vector WE of the character, the position vector PE and the structural subtitle vector LE are added as the input w of the semantic encoding:

[0057]

[0058] 2) Text semantic encoding

[0059] The quality of sentence semantic encoding will significantly affect the performance of extractive summarization. With the development of pre-trained models in recent years, the application of pre-trained models in extractive summarization has brought great improvement to this task. Therefore, it is considered to introduce pre-trained language models into the model of the present invention. However, most of the pre-trained models are currently built based on the Transformer model, which brings huge computational overhead while bringing excellent performance and cannot be directly used in long texts. Previously, some researchers tried to directly truncate the input text to complete the task. Doing so will undoubtedly lose a lot of semantic information of the original document, thereby affecting the performance of the extractive summary. Here, the present invention introduces a pre-trained language model Longformer designed for long texts to adapt to the ultra-long texts in the task.

[0060] Longformer can be regarded as an improved version of the Roberta model for long texts. The main difference between it and the latter is the calculation of the attention mechanism. Unlike the sequence full self-attention of the Transformer structure, the attention mechanism of Longformer consists of two parts, namely local attention and global attention. Each word in the sequence will perform these two attention mechanism calculations. First, each word will perform a complete self-attention calculation on the words of the preset window size, which is the local attention. After that, there are some special characters in the sequence, and each word in the sequence will perform attention calculations on them. It will also perform attention calculations on each word in the sequence, which is the global attention. Finally, the results of these two attentions are combined to obtain the final representation of the word. It can be found that local attention can effectively reduce the amount of calculation of the model, and the existence of global attention allows each word to understand the semantic information of the entire text.

[0061] Like most researchers before, we insert [CLS] characters at the beginning of each sentence and [SEP] characters at the end of each sentence, and then set [CLS] as the special character for global attention in Longformer. Then we input all characters into Longformer to get its hidden layer representation:

[0062] [u 11 ,u12 ,...,u NM ]=Longformer([w 11 ,w 12 ,...,w NM ])

[0063] However, after obtaining its hidden layer representation, the present invention does not use the hidden layer representation of [CLS] as the initialization of sentence nodes like other researchers, because the present invention believes that each word has a different influence on the sentence vector, and it is difficult to get a good effect by directly using [CLS] representation. Therefore, a weighted attention average mechanism is introduced here to obtain the hidden layer representation of sentence nodes:

[0064]

[0065]

[0066]

[0067] Among them, a ij represents the attention weight, u ij Represents the sentence vector representation, W u represents the learnable parameters, and k represents the subscript. Finally, h i As the initialization representation of the sentence.

[0068] 3) Information transmission process

[0069] According to the previously defined hierarchical structure, the obtained sentence vector [h1,h2,...,h N ] as the lowest level initialization vector It also serves as the beginning of information layering transmission. The hidden layer representation of Use Bi-LSTM as a layer encoder to capture the dependencies between them:

[0070]

[0071] here represents the hidden state of the kth sub-segment in the jth paragraph in the i-1th layer after encoding, C i Represents the initial state of the i-th layer BiLSTM.

[0072] Then, if Figure 2 As shown, these updated and enhanced hidden layer representations are averaged and pooled (AVG Pooling) and a forward neural network (MLP) is used to obtain the corresponding hidden layer representation of the i-th layer and the j-th paragraph:

[0073]

[0074] in is the hidden vector representation of the jth paragraph at the i-th level, W s and b S is a learnable parameter, and Avg is mean pooling.

[0075] After the above process, the hidden layer representations from low to high levels are modeled into a sequence, and the information is transmitted from bottom to top according to the encoders of each layer until the vector representation H of the document is obtained. D .

[0076] 4) Information Representation Iteration Unit

[0077] Referring to the process of the previous step, the model can actually obtain a complete hierarchical information representation, but for long texts, this single one-way transmission mechanism is not enough to model the complex semantic information of the model. Therefore, the present invention introduces an iterative enhancement method here, which transfers the text information from the document level back to the sentence level, and performs another round of information transmission, thereby performing iterative updates. To realize the above idea, an attention mechanism is introduced. For the i-th level encoder of the k-th iteration, the hidden layer vector corresponding to the i-th layer of the k-1th round is used to represent H i(k-1) As the query value, the hidden vector corresponding to the i-1th layer obtained in round k-1 represents H i-1(k-1) As the key, calculate the initial state vector C of BiLSTM i (k) The following is a detailed description of the i-th layer encoder in the k-th round:

[0078]

[0079]

[0080]

[0081]

[0082] Where W a , W b , W c and W d are all learnable parameters, is the number of sub-segments to be aggregated at layer i-1 corresponding to the mth sub-segment at layer i. Other values ​​are intermediate calculated values ​​or their meanings have been explained above.

[0083] 5) Information fusion selection

[0084] After layered information transmission and iterative update enhancement, accurate representation of information at each level can be obtained. Then, when scoring the summary sentence, the information of all levels will be integrated to comprehensively evaluate each sentence:

[0085]

[0086] here, is the probability score of the sentence as a summary, W p , W f 、b p and b f is a learnable parameter, : is a vector concatenation operation, H 0 :H 1 :...:H D is the representation of each level, and D is the number of iterations.

[0087] 6) Training objectives

[0088] In order to enable the model to better obtain the information contained in each level structure and each level text, an auxiliary training goal of paragraph attribution is designed for each sentence: the substructure titles with similar meanings in the text are classified into one category, so as to predict which category each sentence belongs to. The probability calculation formula of sentence affiliation category is as follows:

[0089]

[0090] Here W s and b s is a learnable parameter, is the probability distribution of the predicted sentence category.

[0091] Then, the summary sentence selection task and sentence category prediction task are combined as the training objectives of the entire model:

[0092]

[0093] Here ∈1 and ∈2 are hyperparameters for balancing the two training objectives, and CE is the cross entropy loss function.

[0094] The specific steps in the invention content are described in depth below through an embodiment. First, the data set, experimental settings, baseline model and automatic evaluation indicators used in the experiment are introduced.

[0095] 1) Dataset

[0096] The effect of the model on long text summaries is evaluated by using two widely used datasets, arXiv and PubMed, with the same division and preprocessing details. Both datasets are scientific research datasets, and the text structures are similar. They are uniformly modeled as a "sentence-chapter-document" structure for experiments. In order to complete the training of the category prediction target, the chapters of the two datasets are classified separately, with 10 chapter categories in the arXiv dataset and 8 chapter categories in the PubMed dataset.

[0097] 2) Experimental Setup and Baseline Model

[0098] Longformer is used as the semantic encoder of the model, and the specific model parameters come from longformer-base in huggingface / trainsformers. For the BiLSTM used in the sentence-level encoder and the chapter-level encoder, the hidden state vector dimension is set to 768, and each encoder uses two layers of BiLSTM. The model is trained on 3 GPUs (Nvidia Tesla V100, 32G) for 150,000 steps, and the gradient accumulation calculation is performed every two steps. The optimizer selects the Adam optimizer and sets the learning rate to 5e-4, and adds a warmup of 10,000 steps. Here, the parameters of the Longformer pre-trained model are frozen, and only the parameters of the rest are trained. To balance the relationship between the various loss functions, and are set to 0.8 and 0.2 respectively. For both datasets, 7 sentences are selected as summaries.

[0099] In order to further evaluate the model effect, some common extractive summarization models were selected as baseline models for comparison, including:

[0100] Oracle: Based on the greedy algorithm, the extracted sentences are compared with the manual summary, and the combination with the highest ROUGE value is selected as the summary, which can represent the upper limit of the extractive summary.

[0101] LEAD: Directly select the first few sentences of the original document as the summary, because many articles will write important information at the beginning of the document.

[0102] SummaRuNNer: The extractive summarization model based on the recurrent neural network structure proposed by Nallapati et al. comprehensively considers the importance and novelty of the selected sentences, laying a solid foundation for the subsequent extractive summarization model based on deep learning

[0103] Seq2seq-local&global: An extractive summarization model for long texts proposed by Xiao et al., which comprehensively considers the local and global information of the text

[0104] BertSum: An extractive summarization model proposed by Liu et al., which first introduced a pre-trained language model to this task and achieved a significant improvement in performance

[0105] LongSum: Replace the pre-trained language model in BertSum with Longformer

[0106] MatchSum: A two-stage extractive summarization model proposed by Zhong et al., which transforms extractive summarization into a text matching problem

[0107] SSN-DM: An extractive summarization algorithm based on graph neural network and memory network mechanism proposed by Cui et al., which uses the pre-trained model for long texts in the form of sliding windows and achieves good results.

[0108] 3) Automatic evaluation indicators

[0109] The present invention uses the automatic evaluation index ROUGE to judge the results of abstract generation. The open source script ROUGE-1.5.5.pl is used to analyze the results generated by the model and the manual abstracts, and the scores of ROUGE-1, ROUGE-2, and ROUGE-L are combined to evaluate the importance and fluency of abstract generation.

[0110] 4) Positive effects

[0111] 4.1) ROUGE indicator

[0112] The ROUGE scores of the models on the two public datasets are shown in Table 1. The table shows the effects of the basic model (Oracle and Lead), the traditional deep learning extractive model and the model from top to bottom. Obviously, the model exceeds all baseline models in the R-1 and RL evaluation indicators, which shows that making full use of the structural information in the text and modeling the text multiple times can significantly improve the quality of summary selection. Xiao et al. and Nallapati et al. are both extractive summary models based on recurrent neural network structures. The scores of the latter are better than the former, indicating that their comprehensive modeling of global and local information has a significant effect on the long text summary task. This is similar to the ideas of hierarchical modeling and multiple modeling, and also indirectly confirms the correctness of the ideas of the present invention. The results of the LongSum model exceed those of these models using traditional static word vectors, indicating that the introduction of the pre-training model has a great improvement on the summary task. The model of the present invention combines the advantages of these two types of models and is designed specifically for the characteristics of these two scientific research datasets, thereby achieving the optimal results.

[0113] Table 1 Automatic evaluation index results of the extractive summarization model based on hierarchical iteration

[0114]

[0115] 4.2) Module ablation experiment

[0116] In order to explore the roles played by many modules in the model, an ablation experiment was conducted on the arXiv dataset, and the experimental results are shown in Table 2. This ablation experiment includes the following ablation modules:

[0117] a) w / o hierarchical position: the hierarchical position vector corresponding to the character is not added during input, and only the word sequence position information is retained.

[0118] b) w / o section information, the text vector of the chapter is not added during input, and the chapter classification auxiliary task is deleted.

[0119] c) w / o sentence encoder, delete the sentence-level encoder, directly use the sentence vector to get the document vector, and use the sentence vector and chapter vector for prediction.

[0120] d) w / o section encoder, delete the chapter-level encoder, no longer generate document vectors, and use sentence vectors and chapter vectors for prediction.

[0121] e) w / o both encoders, two levels of encoders are deleted, no information is transferred, and only sentence vector prediction is used.

[0122] Table 2 Ablation experiments of HISum on arXiv

[0123] Model R-1 R-2 RL HISum 45.22 17.67 40.02 HISum w / o hierarchical position 45.18 17.65 39.99 HISum w / o section information 44.53 17.49 39.71 HISum w / o section encoder 44.37 17.44 39.84 HISum w / o document encoder 44.86 17.21 39.87 HISum w / o both encoder 43.92 17.12 39.75

[0124] From the results of the ablation experiment, we can find that the module with the least impact on the model effect is the hierarchical position encoding, that is, adding the position vector of the sentence and chapter where the input character is located does not play a big role, which shows that the model is not very sensitive to the position information of the sentence, and it is far from enough to use only the position information to represent the structural information of the document; adding chapter information and introducing auxiliary tasks have greatly improved the model in ROUGE value, which shows that the information contained in the chapter itself is of great help to the selection of the summary, and the chapter title can represent the core content of the chapter to a certain extent; the sentence-level encoder and the chapter-level encoder have a positive impact on the model, and they can achieve the best effect when they exist together in the model, which also proves the effectiveness of the hierarchical transfer method in modeling text structure information.

[0125] 4.3) Number of iterations

[0126] Finally, the influence of the number of iterations on the model effect was explored. The experimental results are as follows: Figure 3 As shown. Figure 3 From the curve changes in , we can find that when the number of iterations is small, the effect of the model increases with the increase of the number of iterations, which shows that the idea of ​​multiple modeling text information is correct in the summary task. However, as the number of iterations increases, the model effect begins to decrease to a certain extent. This may be because when the number of iterations is too large, the information contained in each level of the model begins to become vague, and each sentence may contain the central information of the chapter or document to a certain extent, thereby interfering with the selection of the summary. Therefore, after many experiments, the optimal number of iterations is set to 3, which can ensure the superior performance of the model and reduce unnecessary computing resources.

[0127] In summary, the present invention proposes a long text extractive summarization method based on hierarchical iteration, which models the long text as a hierarchical sequence structure, hierarchically models the complex text structure of the long text to complete the hierarchical transmission of information, and then iteratively updates the three-level representation vectors to model the text multiple times to improve the model understanding effect. The results on the arXiv and PubMed datasets show that the model of the present invention is superior to the previous baseline model. Subsequently, a more detailed experiment was designed for the results of the model for further analysis, which verified the rationality of the model design of the present invention and the effectiveness of the parameter setting.

[0128] Another embodiment of the present invention provides a long text extractive summary generation device based on hierarchical iteration, which includes:

[0129] The preprocessing module is used to obtain the word vector, position vector and structure subtitle vector of the characters in the text;

[0130] The text semantic encoding module is used to add the word vector, position vector and structure subtitle vector as the input of semantic encoding, and use the long text pre-trained language model as the semantic encoder to perform semantic encoding;

[0131] The information transfer module is used to send the semantically encoded vectors to the encoders at each level, transfer the semantic information hierarchically from the sentence level to the document level along the text structure route, and then transfer it hierarchically from the document level to the sentence level again, to achieve iterative update and obtain the hidden layer representation of each level;

[0132] The information fusion selection module is used to comprehensively evaluate each sentence by fusing the hidden layer representations of each level and select the optimal summary sentence.

[0133] The specific implementation process of each module refers to the above description of the method of the present invention.

[0134] Another embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.

[0135] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk), wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the steps of the method of the present invention are implemented.

[0136] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. A person skilled in the art may modify or make equivalent substitutions for the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A hierarchical iterative long text extractive summary generation method, characterized in that: The following steps are involved: The hierarchical structure of the text is defined, and the minimum hierarchical structure of an article is set as sentence S, and the maximum hierarchical structure is set as article D. In between, there are several layers of structure L = {L 1 ,L 2 ,...,L n }, for the text paragraph at level i have in represents the mth subparagraph in the jth paragraph; Get the word vector, position vector, and structure subtitle vector of the characters in the text; The position vector is obtained by calculating the position vector PE by treating the characters as a sequence. word , treat each level structure as a sequence and calculate the position vector PE L , for those in The characters in have the same i-th level position vector The final position vector PE is obtained by summing the position vectors of each level; The word vector, position vector and structure subtitle vector are added together as the input of semantic encoding, and the long text pre-trained language model is used as the semantic encoder for semantic encoding; The semantically encoded vectors are sent to encoders at each level, and the semantic information is transferred hierarchically from the sentence level to the document level along the text structure route, and then transferred hierarchically from the document level to the sentence level again, to achieve iterative updates and obtain the hidden layer representations of each level; By integrating the hidden representations of each level, each sentence is comprehensively evaluated and the best summary sentence is selected; The long text pre-trained language model is a Longformer model, and the attention mechanism of the Longformer model includes local attention and global attention; each word in the sequence is calculated by two attention mechanisms. First, each word performs a complete self-attention calculation on words of a preset window size, that is, local attention. Then, each word in the sequence performs an attention calculation on special characters, that is, global attention. Finally, the results of the two attention mechanisms are combined to obtain the final representation of the word. The encoders at each level are constructed by BiLSTM. The hidden layer representation of the document is obtained by using Bi-LSTM as a hierarchical encoder to capture the dependencies between them. The hidden layer representation is then mean-pooled and passed through a forward neural network to obtain the corresponding hidden layer representation of the jth paragraph of the i-th layer. The hidden layer representations from low to high levels are all modeled as a sequence. The encoders of each layer transmit the information from bottom to top in layers until the vector representation of the document is obtained. The layered transfer from the document level to the sentence level is performed again to achieve iterative updating, including: introducing an attention mechanism, for the i-th level encoder of the k-th iteration, using the hidden layer vector representation corresponding to the i-th layer in the k-1th round as the query value, and the hidden layer vector representation corresponding to the i-1th layer obtained in the k-1th round as the key value, and calculating the initial state vector of the BiLSTM.

2. The method according to claim 1, characterized in that The Longformer model uses the following steps to get the initial representation of the sentence: Insert the [CLS] character at the beginning of each sentence and the [SEP] character at the end of the sentence, then set [CLS] as the special character for global attention in Longformer, and then input all characters into Longformer to get its hidden layer representation: [u 11 ,u 12 ,...,u NM ]=Longformer([w 11 ,w 12 ,...,w NM ]) A weighted attention average mechanism is introduced to obtain the hidden layer representation of sentence nodes: Among them, a ij represents the attention weight, u ij Represents the sentence vector representation, W u represents the learnable parameters, k represents the subscript, and finally h i As the initialization representation of the sentence.

3. The method according to claim 1, characterized in that During the model training process, an auxiliary training goal of paragraph attribution is designed for each sentence: the substructure titles with similar meanings in the text are classified into one category, so as to predict which category each sentence belongs to; then the summary sentence selection task and the sentence category prediction task are combined as the overall training goal of the model.

4. A hierarchical iterative long text extractive summary generation device using the method according to any one of claims 1 to 3, characterized in that: include: The preprocessing module is used to obtain the word vector, position vector and structure subtitle vector of the characters in the text; The text semantic encoding module is used to add the word vector, position vector and structure subtitle vector as the input of semantic encoding, and use the long text pre-trained language model as the semantic encoder to perform semantic encoding; The information transfer module is used to send the semantically encoded vectors to the encoders at each level, transfer the semantic information hierarchically from the sentence level to the document level along the text structure route, and then transfer it hierarchically from the document level to the sentence level again, to achieve iterative update and obtain the hidden layer representation of each level; The information fusion selection module is used to comprehensively evaluate each sentence by fusing the hidden layer representations of each level and select the optimal summary sentence.

5. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Long text data processing method and device, computing equipment, storage medium and product

    CN116151216A