Method and system for obtaining text timeline summary
Through a unified timeline abstract generation model, combined with generative and extraction methods, the difficulty of time series information capture and abstract authenticity problems in the prior art is solved, and efficient and accurate timeline abstract generation is achieved.
Patent Information
- Application Number
- CN202210803029.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-07-07
AI Technical Summary
The existing timeline digest methods mainly rely on the extracted method, making it difficult to capture time series information and implement highly flexible generative digests, and authenticity problems lead to information confusion and error digests.
A unified timeline digest generation model is proposed, combining generative and decimation methods, events are encoded through bidirectional recursive neural network and selective reading unit, time series information is captured using multi-layer perceptron and key-value memory modules, and a time-order summary is generated through an RNN-based decoder.
It realizes the generation of a timeline summary that correctly summarizes the information related to the original text, improves the authenticity and accuracy of the abstract, and can output the generated and extracted abstracts in chronological order, which promotes collaborative learning between the two tasks.
Smart Images

Figure CN115221312B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method and system for obtaining a timeline summary of a text, namely a generative summary and an extractive summary. Background Art
[0002] In this era of information explosion, readers are drowned in a sea of text information and do not know where to get concise and effective information. General search engines only return web pages sorted by the relevance of query keywords, but cannot handle queries with unclear intentions, such as queries about a certain type of news. In fact, people may be particularly concerned about the beginning, evolution or latest developments of an event, while ordinary information retrieval technology only sorts according to the relevance of the search terms and then returns the sorted web page results.
[0003] In addition, in many cases, even if the query results can be arranged in a satisfactory order, readers are tired of browsing a large number of documents. They prefer to understand the evolution of current affairs and hot events through simple and fast browsing. For this problem, summary is a suitable solution. It can provide a brief document summary containing important information, and summarize real-time news faster and better. Furthermore, the timeline summary summarizes the real-time news into a series of independent but related summary combinations, so that readers can understand the development of the event as a whole. Among them, the extractive summary directly extracts sentences from the original text, which can produce fluent sentences. The generative summary has higher flexibility and can generate phrases and phrases that are not in the original text.
[0004] Existing timeline summarization methods are all based on extractive methods. However, these methods rely on complex features that are manually defined. Some other model features that are very important for summarization, such as the ability to explain, summarize, or introduce other knowledge, can only be achieved in a generative framework. Recently, as models that use deep neural networks for text generation have received increasing attention, generative techniques have become increasingly popular. Generative summarization methods have been proven to be very useful in traditional summarization tasks. However, unlike traditional document summarization, timeline summary datasets consist of a series of events with timestamps, so timeline summary models must capture this time series information to better guide the generation process of timeline summarization. In addition, the authenticity issue is also an important issue for timeline summarization. The mixing of information from different events often leads to incorrect summaries. Summary of the invention
[0005] The purpose of the present invention is to propose a method and system for obtaining a timeline summary of a text, and to construct a unified timeline summary generation model, which can output two timeline summaries, namely, a generative summary and an extractive summary, which can correctly summarize the relevant information in the original text in chronological order.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides a method for obtaining a text timeline summary, comprising the following steps:
[0008] Taking an original document as input, the original document includes events at different time nodes along a timeline, and each event includes a group of words;
[0009] The encoding matrix is used to obtain the encoding representation of each word in each event, and the interaction between words is modeled through the LSTM of the bidirectional recurrent neural network to obtain the word representation; the word representation is then processed through the selective read unit SRU to obtain the local event representation.
[0010] Multi-layer perceptron (MLP) is used to establish relationship edges between two event representations in the document modeling graph, and the relationship edges are incorporated into the attention mechanism operation in the relationship-aware encoding process to obtain a global event representation.
[0011] Construct a key-value memory module, wherein the key part of the key-value memory module stores time points through time position encoding; the value part stores local event representations through local values and stores global event representations through global values; and extracts information from the value part using the key as guidance information of time;
[0012] Using an RNN-based decoder, the initial state of the decoder is obtained by randomly initializing the LSTM unit, and the association weight between the decoder state and each word representation is calculated to obtain the attention distribution; the word context vector is calculated based on the attention distribution;
[0013] Calculate the weighted sum of local event representations to obtain the event context vector;
[0014] According to the correlation between the time position encoding in the key-value memory module and the current state of the decoder, the local value in the key-value memory module is changed to the memory model vector; a summary is generated according to the decoder state, the word context vector, the event context vector and the memory model vector;
[0015] Use a bidirectional RNN to process each sentence in the input document to obtain sentence representation and initialized document representation; use an SRU-based RNN to iteratively update the sentence representation and document representation to obtain better sentence representation and better document representation;
[0016] Through the LSTM-based RNN, the attention weights of the input sentences are calculated in chronological order, and the sentence context vector is obtained based on the weights; the index of the sentence is obtained based on the sentence context vector, and the index is used to extract the sentences in chronological order as summaries.
[0017] Furthermore, SRU obtains local event representations based on word representations, coarse-grained event representations, and the hidden states of SRU units.
[0018] Furthermore, the method of obtaining the initial state of the decoder by randomly initializing the LSTM unit is as follows: by randomly initializing an LSTM unit, using the unit to take the concatenated representation of all local event representations as input, and taking the input as the initial state of the decoder.
[0019] Furthermore, a method for calculating a word context vector according to the attention distribution is: using the attention distribution to obtain a weighted sum of document representations to obtain a word context vector.
[0020] Furthermore, according to the correlation between the time position code in the key-value memory module and the current state of the decoder, the local value in the key-value memory module is changed into a memory model vector. The specific steps include: first, using the current decoder state to read each key in the time-event memory module to obtain the time position code; second, calculating the correlation between the time position code and the current decoder state, and using it as the key of the time attention weight; third, according to the key of the time attention weight, through the fusion gate, the local value in the time-event memory module is changed into a memory model vector and merged into the projection layer.
[0021] Furthermore, a method for generating a summary based on the decoder state, the word context vector, the event context vector, and the memory model vector is as follows: the outputs of the decoder state, the word context vector, the event context vector, and the memory model vector are connected and input into the projection layer to obtain a final generated distribution on the vocabulary, which is the summary.
[0022] Furthermore, a bidirectional RNN is used to process each sentence of the input document, and a method for obtaining a sentence representation and an initialized document representation is as follows: a bidirectional RNN is used to process each sentence of the input document, and the last hidden state is used to represent the entire sentence representation; and an average value of all sentence representations is calculated and used as the initialized document representation.
[0023] Furthermore, the SRU-based RNN is used to introduce the hidden state of the RNN to iteratively update the sentence representation and document representation.
[0024] Furthermore, a method for obtaining a sentence index based on the sentence context vector is as follows: combining the sentence context vector and the sentence representation through LSTM to update the hidden state; and processing the hidden state through a multi-layer perceptron MLP and an argmax function to obtain the sentence index.
[0025] The present invention also provides a system for obtaining a text timeline summary, which includes an event encoding module, a graph-based encoder, a time-event memory module and a summary generator for performing a generative summary task, a sentence encoding module and a summary extractor for performing an extractive summary task, and an attention unification module for penalizing the inconsistency between the generative summary task and the extractive summary task according to a time-aware inconsistency loss function, wherein:
[0026] The event encoding module uses the encoding matrix to obtain the encoding representation of each word in each event; the interaction between words is modeled through the LSTM of the bidirectional recurrent neural network to obtain the word representation; the local event representation is obtained according to the word representation, the coarse-grained event representation and the hidden state of the SRU unit through the selective reading module;
[0027] The graph-based encoder uses a multi-layer perceptron (MLP) to establish relationship edges between two event representations in the document modeling graph, and incorporates the relationship edges into the attention mechanism operation during the relationship-aware encoding process to obtain a global event representation.
[0028] The time-event memory module is a key-value memory module, in which the key part stores time points through time position encoding; the value part stores local event representations through local values and global event representations through global values; the key is used as the guiding information of time to extract information from the value part;
[0029] The summary generator is a RNN-based decoder. It randomly initializes an LSTM unit, takes the concatenated representation of all local event representations as the input of the decoder, and uses the input as the initial state of the decoder; uses the previous decoder state to calculate its association weight with each word representation to obtain the attention distribution; uses the attention distribution to obtain the weighted sum of the document representation to obtain the word context vector; calculates the weighted sum of the local event representation to obtain the event context vector; uses the current decoder state to read each key in the time-event memory module to obtain the time position code; calculates the correlation between the time position code and the current decoder state and uses it as the key of the time attention weight; according to the key of the time attention weight, changes the local value in the time-event memory module to the memory model vector through the fusion gate and merges it into the projection layer; connects the outputs of the decoder state, word context vector, event context vector and memory model vector, and inputs them into the projection layer to generate a summary;
[0030] The sentence encoding module uses a bidirectional RNN to process each sentence of the input document, and uses the last hidden state to represent the entire sentence representation; calculates the average of all sentence representations and uses it as the initialized document representation; uses the SRU-based RNN, introduces the hidden state of the RNN, iteratively updates the sentence representation and document representation, and obtains better sentence representation and better document representation;
[0031] The summary extractor is an LSTM-based RNN. It calculates the attention weights of the input sentences in chronological order and obtains the sentence context vector based on the weights. It updates the hidden state by combining the sentence context vector with LSTM and obtains the index of the selected sentence by mapping the hidden state. It uses the index to extract sentences as summaries in chronological order.
[0032] The technical solution of this invention actually includes two parts. The first part is the generative summary part. On the encoder side, it is responsible for reading event documents containing multiple time nodes, obtaining document representations at the word level and event level, and modeling the relationship between multiple input events; on the decoder side, it aims to summarize the main idea of the input document with new words and phrases, and output a generative summary. The second part is the extractive summary part, which uses an iterative method to obtain better article and sentence representations, and finally obtains an extractive summary. The present invention unifies these two parts, can output generative and extractive summaries in chronological order, and link the two tasks to promote each other.
[0033] In the encoder part, the present invention proposes a graph-based event encoding module, which learns the correlation of multiple events according to the dependencies between the contents of multiple events, and further learns the global representation of each event. In the decoding part, in order to ensure the temporal order of the generative summary, the present invention proposes to extract the characteristics of event-level attention in the summary generation process while retaining the sequence information, and use it to simulate the attention evolution process of the standard summary. The event-level attention in the generative summary can also be used to assist the temporal order of the extractive summary. The attention unification module proposes a time-aware inconsistency loss function to penalize the inconsistency between the two tasks. Whether it is the generative summary task or the extractive summary task, the attention to the input document should follow the temporal order. Therefore, it is intuitive to encourage the two levels of attention to be basically consistent during the training process and use this as an intrinsic learning goal without additional cost. The present invention first uses a convolutional neural network (CNN) to extract the evolving event attention features from the summarizer and extractor, and obtains a new attention matrix whose abscissa length is the number of extracted sentences by convolution along the horizontal axis on the event-level attention map. Afterwards, for each sentence extraction step, the event-level attention is replicated multiple times. The present invention hopes that when the attention at the sentence level is high, the attention at the event level is also high. Therefore, the present invention designs a time-aware inconsistency loss function. This loss function encourages the attention at the event level corresponding to the sentence with high attention to be correspondingly higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of a system for obtaining a text timeline summary proposed by the present invention;
[0035] Figure 2 A schematic diagram of a generative summary part proposed by the present invention;
[0036] Figure 3 A schematic diagram of the extractable summary part proposed by the present invention;
[0037] Figure 4 A schematic diagram of the attention unification module proposed in the present invention; DETAILED DESCRIPTION
[0038] In order to make the above features and advantages of the present invention more obvious and easy to understand, embodiments are given below and described in detail with reference to the accompanying drawings.
[0039] This example uses a timeline summary as an example. As shown in the following text, the incorrect summary confuses birthplace and residence, first album and best-selling album. Preliminary experimental results show that this phenomenon of wrong information is a common problem in summary tasks.
[0040]
[0041] A good timeline summary correctly summarizes the relevant information in the original article, as shown in the following text:
[0042]
[0043] In the timeline summary task scenario studied, this embodiment provides a system for obtaining a text timeline summary, such as Figure 1 As shown in Figure 1, the system is a timeline-based generative and extractive summarization system. The system includes an event encoding module for performing generative summarization tasks, a graph-based encoder, a time-event memory module, and a summary generator (see Figure 2 ), a sentence encoding module and a summary extractor for performing extractive summarization tasks (see Figure 3 ), and an attention unification module that penalizes inconsistencies between the generative and extractive summarization tasks according to a time-aware inconsistency loss function (see Figure 4 ). In the generative summary part, a graph-based encoder is used to learn the correlation of multiple events according to the dependencies between the contents of multiple events, and further learn the global representation of each event. In the decoding part, in order to ensure the temporal order of the generative summary, this embodiment proposes to extract the features of event-level attention in the summary generation process while retaining the sequence information, and use it to simulate the evolution process of attention of the standard summary.
[0044] This system takes the original document as input, which can be represented as a list of events. Where T e is the number of events. Each event x i is a list of words: in is the jth word in the i-th event, is event x i The number of words.
[0045] This embodiment aims to generate a summary for the generative summary part. $ is the summary symbol, is the i-th word generated, where i = 1,…,T y , T y is the number of words used to generate the summary. The summary must not only be grammatically correct, but also consistent with the event information (such as the location and time of occurrence). In essence, this embodiment attempts to optimize the parameters to maximize its probability in is a standard timeline summary, t represents a single word in Y, where t = 1, ..., T y .
[0046] For the extractive summarization part, the goal is to generate a score vector for each sentence represents the probability of the i-th sentence being extracted, where i=1,…,T ys , T ys is the number of sentences. This embodiment converts the standard summary written by human experts into a standard vector Among them, l i Indicates whether the i-th sentence is selected, where i = 1,…,T ys , if selected, it is equal to 1, and if not selected, it is equal to 0. During the training process, L and Calculate the cross entropy loss and use it to optimize the model.
[0047] The specific description of each module included in this system is as follows:
[0048] Event encoding module:
[0049] First, this embodiment uses the encoding matrix e to convert x i The representation of each word in is mapped to a high-dimensional vector space. Representing words After that, this embodiment uses the LSTM (Long Short-Term Memory Network) of the bidirectional recurrent neural network (BI-RNN) to model the interaction between words (the arrow indicates the recursive direction):
[0050]
[0051] In addition to obtaining word representation It is also necessary to obtain an event representation. If the final state of BI-RNN is simply used as the representation of the entire event, the characteristics of the entire event cannot be fully captured. Therefore, this embodiment uses a selective reading module composed of SRU to obtain a new event representation, namely a local event representation a i .
[0052]
[0053] in, is the hidden state of the t-th SRU unit in the ith event, is the representation of the Teth moment. Overall, SRU is an improved version of the gated recurrent unit GRU, which replaces the update gate in the original GRU with a new gate that considers each input and coarse-grained event representation The SRU unit receives two inputs, the word representation and the sentence representation from the previous iteration.
[0054] Graph-based encoders:
[0055] The local event in the previous section represents a i is calculated independently without considering the information flow between different events. However, the importance of each event and whether it should be adopted as a summary depends not only on itself but also on the content of other events. For example, Ethan Hope's release of his first solo album may be an important event, but his "Firefly" album broke the ideological barrier and was more important than the release of his first solo album. Therefore, this event weakened the importance of the first event. Based on such considerations, this embodiment proposes a graph-based encoder to learn the relationship between events and obtain a global event representation containing this information.
[0056] The encoder builds relationship edges in the document modeling graph to encode relationship information. The relationship edges in the graph are first represented by connection events and initialized through the MLP layer:
[0057] r i,j =MLP(a i ; a j )
[0058] Where i and j are the event numbers, and MLP is a multi-layer perceptron. Next, in the relation-aware encoding process (RE), the relation edge r i,j Merge into the attention mechanism operation:
[0059] b i =RE(a i ,a * ,r i,* )
[0060] Where * represents any index number. The Transformer architecture is based on the fact that each event is not isolated and its representation also depends on other input events. In the enhanced relation-aware coding of this embodiment, the updated event representation is the global event representation b i It also follows the same idea and contains information from other documents. The difference is that b i Not only depends on other documents, but also on b i In other words, the representation of each event consists of its dependencies with other events, thus representing the attention of the event more comprehensively.
[0061] Time-Event Memory Module:
[0062] As mentioned above, in the timeline summarization task, the generated summary should capture the time series information. Therefore, this embodiment proposes a time-event memory module, which is a key-value memory module, where all the keys together form a timeline for guiding the summary generation process. The key in this memory module is the time position code introduced in the event encoding module. i This embodiment will use this key as the guidance information of time to extract information from the value part in the memory model. The value part in the memory module includes local event representation and global event representation. The local value only stores the local event representation a i , which means that it only captures the local information of the current event. The global value only stores the global event representation and is responsible for learning the characteristics of the event from a global perspective. This feature is not only based on itself, but also based on its relationship with other events. Therefore, it is responsible for storing the output of the graph-based encoder, that is, the global event representation b i .
[0063] Summary Generator:
[0064] In order to generate a coherent and consistent summary, this embodiment proposes a summary generator, which is a RNN-based decoder that reads the global event representation of the time-event memory module and the encoder and uses it as input. First, this embodiment randomly initializes an LSTM unit, which takes the concatenated representation of all events (that is, the concatenated representation of all local event representations [a 1 ,…,a T ] as input, and use that input as the decoder initial state:
[0065] h′ 0 =LSTM(h c ,[a 1 ,…,a T ])
[0066] Where T is the total number of time steps and h is the decoder state. Next, following the traditional attention mechanism, this embodiment dynamically aggregates the input document information into the word context vector c t-1 Here, this embodiment first uses the decoder state h′ t-1 To calculate it with each state The associated weights of in, Represents event x i The representation of the jth word in . Afterwards, this embodiment uses the attention distribution To obtain the weighted sum of the document representation, as the word context vector c t-1 .
[0067] In addition to using event-level attention to directly guide word-level attention, this embodiment also uses event-level attention to obtain the weighted sum of local event representations, namely the event context vector e t , in the following formula is the weight of the local vector representation:
[0068]
[0069] Next, we will introduce how to use the guidance information from the memory model. This embodiment first uses the current decoder state (hidden state) h′ t to read each key in the time-event memory module. The key in the previous section, i.e., the temporal position code, constitutes a timeline representing the temporal sequence information. Therefore, this embodiment allows the model to use this sequence information to calculate the correlation between the temporal position code and the current decoder state as the key π(p i ,h′ t ), used to obtain where p i , is the temporal position code, h′ t is the decoder state. Through the fusion gate, the local value is changed to and will be merged into the projection layer, is the memory model vector output by the time-event memory module. The reason why the local value is placed in the projection layer in this embodiment is that It stores local information rather than global features. Therefore, it should work when generating each specific word. It is responsible for storing global characteristics of events and should affect the entire generation process at a holistic level. The information is fused into the decoding state h′ through a gate t .
[0070] Finally, the summarizer passes through an output projection layer to obtain the final generated distribution over the vocabulary:
[0071]
[0072] In this embodiment, the decoder state h′ t , word context vector c t , event context vector e t and memory model vector The outputs of are connected and used as the input of the output projection layer. To solve the vocabulary problem, this embodiment adapts the point / copy mechanism on the decoder so that the decoder can copy words from the source text.
[0073] Sentence encoding module:
[0074] This example first uses a new bidirectional RNN to process each sentence and obtains the representation Representing the tth word in the ith sentence uses the last hidden state to represent the entire sentence representation. The document representation is initialized as the average of all sentence representations:
[0075]
[0076] In the formula, W and b are two constant coefficients.
[0077] Next, in order to establish a sequential relationship model between sentences and obtain a more comprehensive sentence representation, this embodiment iteratively polishes the sentence and document representation. For the sake of brevity, this embodiment uses the first iteration as an example to illustrate this process. Specifically, there is an SRU-based RNN in the iterative process:
[0078]
[0079] In the formula, D is the document representation, a is the sentence representation, is the hidden state of the RNN.
[0080] Get better sentence representation through this RNN Based on this sentence representation, we can get a better document representation D 2 . The document then indicates D 2 Alternative D 1 , once again improved Update it to In this way, this embodiment uses a loop iteration method to improve the sentence and document representation. This embodiment uses I to represent the number of iterations, so the final representation of the Ith sentence is
[0081] Abstract Extractor:
[0082] The summary extractor follows the traditional attention mechanism. The module calculates the sentence-level attention weights for all input sentences, obtains the context vector based on this weight, and updates the hidden state in combination with the context vector. Finally, the index of the selected sentence is obtained by mapping the hidden layer state. In this way, the summary extractor extracts sentences as summaries in chronological order to better capture the sequential information in the input. This embodiment uses an RNN composed of LSTM units to select sentences. Each decoding step extracts an important sentence from the original text in chronological order as a summary. Following the traditional attention mechanism, the module calculates the sentence-level attention weights for all input sentences, obtains the sentence context vector based on this weight, and updates the hidden state in combination with the context vector. Finally, the index of the selected sentence is obtained by mapping the hidden layer state. In this way, the summary extractor extracts sentences as summaries in chronological order to better capture the sequential information in the input. And update the hidden state based on the context vector Finally, through the mapping The index ot of the selected sentence is obtained, where t here represents the sentence selected at the t-th step. Specifically, the t-th decoding step is calculated as follows:
[0083]
[0084]
[0085] in, is the weight obtained at the tth step of the i-th sentence, and I is the number of iterations. In this way, the summary extractor extracts sentences as summaries in chronological order to better capture the sequential information in the input.
[0086] Attention Unification Module:
[0087] like Figure 4 As shown in Figure 2, the module penalizes the inconsistency between the generative and extractive summarization tasks according to the time-aware inconsistency loss function, encouraging sentences with high attention to have correspondingly higher attention at the event level.
[0088] Although the present invention has been disclosed as above by way of embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent substitutions of the technical solutions of the present invention made by ordinary technicians in the field should all be included in the protection scope of the present invention. The protection scope of the present invention shall be based on what is defined in the claims.
Claims
1. A method for obtaining a text timeline summary, It is characterized in that The following steps are involved: Taking an original document as input, the original document includes events at different time nodes along a timeline, and each event includes a group of words; The encoding matrix is used to obtain the encoding representation of each word in each event, and the interaction between words is modeled through the LSTM of the bidirectional recurrent neural network to obtain the word representation; the word representation is then processed through the selective read unit SRU to obtain the local event representation. Multi-layer perceptron (MLP) is used to establish relationship edges between two event representations in the document modeling graph, and the relationship edges are incorporated into the attention mechanism operation in the relationship-aware encoding process to obtain a global event representation. Construct a key-value memory module, wherein the key part of the key-value memory module stores time points through time position encoding; the value part stores local event representations through local values and stores global event representations through global values; and extracts information from the value part using the key as guidance information of time; Using an RNN-based decoder, the initial state of the decoder is obtained by randomly initializing the LSTM unit, and the association weight between the decoder state and each word representation is calculated to obtain the attention distribution; the word context vector is calculated based on the attention distribution; Calculate the weighted sum of local event representations to obtain the event context vector; According to the correlation between the time position encoding in the key-value memory module and the current state of the decoder, the local value in the key-value memory module is changed to the memory model vector; a summary is generated according to the decoder state, the word context vector, the event context vector and the memory model vector; Use a bidirectional RNN to process each sentence in the input document to obtain sentence representation and initialized document representation; use an SRU-based RNN to iteratively update the sentence representation and document representation to obtain better sentence representation and better document representation; Through the LSTM-based RNN, the attention weights of the input sentences are calculated in chronological order, and the sentence context vector is obtained based on the weights; the index of the sentence is obtained based on the sentence context vector, and the index is used to extract the sentences in chronological order as summaries.
2. The method according to claim 1, It is characterized in that SRU obtains local event representation based on word representation, coarse-grained event representation and hidden state of SRU unit.
3. The method according to claim 1, It is characterized in that The method of obtaining the initial state of the decoder by randomly initializing an LSTM unit is as follows: by randomly initializing an LSTM unit, using the concatenated representation of all local event representations of the unit as input, and using the input as the initial state of the decoder.
4. The method according to claim 1, It is characterized in that The method for calculating the word context vector based on the attention distribution is: using the attention distribution to obtain the weighted sum of the document representation to obtain the word context vector.
5. The method according to claim 1, It is characterized in that According to the correlation between the time position encoding in the key-value memory module and the current state of the decoder, the local value in the key-value memory module is changed to a memory model vector. The specific steps include: first, using the current decoder state to read each key in the time-event memory module to obtain the time position encoding; second, calculating the correlation between the time position encoding and the current decoder state, and using it as the key of the time attention weight; third, according to the key of the time attention weight, through the fusion gate, the local value in the time-event memory module is changed to a memory model vector and merged into the projection layer.
6. The method according to claim 1, It is characterized in that The method for generating a summary based on the decoder state, word context vector, event context vector and memory model vector is: connect the outputs of the decoder state, word context vector, event context vector and memory model vector, and input them into the projection layer to obtain the final generated distribution on the vocabulary, which is the summary.
7. The method according to claim 1, It is characterized in that The method of using a bidirectional RNN to process each sentence of the input document and obtaining a sentence representation and an initialized document representation is as follows: using a bidirectional RNN to process each sentence of the input document, using the last hidden state to represent the entire sentence representation; calculating the average of all sentence representations and using it as the initialized document representation.
8. The method according to claim 1, It is characterized in that Using the SRU-based RNN, the hidden state of the RNN is introduced to iteratively update the sentence representation and document representation.
9. The method according to claim 1, It is characterized in that The method for obtaining the index of a sentence based on the sentence context vector is as follows: combining the sentence context vector and the sentence representation through LSTM to update the hidden state; and processing the hidden state through a multi-layer perceptron MLP and an argmax function to obtain the index of the sentence.
10. A system for obtaining a text timeline summary, It is characterized in that It includes an event encoding module, a graph-based encoder, a time-event memory module, and a summary generator for performing a generative summary task, a sentence encoding module and a summary extractor for performing an extractive summary task, and an attention unification module that penalizes the inconsistency between the generative summary task and the extractive summary task according to a time-aware inconsistency loss function, wherein: The event encoding module uses the encoding matrix to obtain the encoding representation of each word in each event; the interaction between words is modeled through the LSTM of the bidirectional recurrent neural network to obtain the word representation; the local event representation is obtained according to the word representation, the coarse-grained event representation and the hidden state of the SRU unit through the selective reading module; The graph-based encoder uses a multi-layer perceptron (MLP) to establish relationship edges between two event representations in the document modeling graph, and incorporates the relationship edges into the attention mechanism operation during the relationship-aware encoding process to obtain a global event representation. The time-event memory module is a key-value memory module, in which the key part stores time points through time position encoding; the value part stores local event representations through local values and global event representations through global values; the key is used as the guiding information of time to extract information from the value part; The summary generator is a RNN-based decoder. It randomly initializes an LSTM unit, takes the concatenated representation of all local event representations as the input of the decoder, and uses the input as the initial state of the decoder; uses the previous decoder state to calculate its association weight with each word representation to obtain the attention distribution; uses the attention distribution to obtain the weighted sum of the document representation to obtain the word context vector; calculates the weighted sum of the local event representation to obtain the event context vector; uses the current decoder state to read each key in the time-event memory module to obtain the time position code; calculates the correlation between the time position code and the current decoder state and uses it as the key of the time attention weight; according to the key of the time attention weight, changes the local value in the time-event memory module to the memory model vector through the fusion gate and merges it into the projection layer; connects the outputs of the decoder state, word context vector, event context vector and memory model vector, and inputs them into the projection layer to generate a summary; The sentence encoding module uses a bidirectional RNN to process each sentence of the input document, and uses the last hidden state to represent the entire sentence representation; calculates the average of all sentence representations and uses it as the initialized document representation; uses the SRU-based RNN, introduces the hidden state of the RNN, iteratively updates the sentence representation and document representation, and obtains better sentence representation and better document representation; The summary extractor is an LSTM-based RNN. It calculates the attention weights of the input sentences in chronological order and obtains the sentence context vector based on the weights. It updates the hidden state by combining the sentence context vector with LSTM and obtains the index of the selected sentence by mapping the hidden state. It uses the index to extract sentences as summaries in chronological order.
Citation Information
Patent Citations
Timeline abstract automatic generation method based on event detection technology
CN113254632A
Automated Timeline Completion Using Event Progression Knowledge Base
US20170351754A1