Event Summarization via Keyword-Entity Sentence Ordering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current risk mining technologies face challenges in efficiently processing large volumes of textual data to identify relevant risk events, as they require significant computational resources and energy, and often produce summaries that are not informative enough for analysts due to the use of abstractive summarization methods which are resource-intensive.
Innovation Solution
The system extracts and orders sentences based on keyword and entity pairs, alternating between specific and general sentences to generate summaries that are more readable and require fewer computational resources, using a combination of user-generated and automatically expanded keyword sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If abstractive summarization techniques are used to generate summaries, then the summaries are preferred by humans for content and readability, but the computational expense and hardware resources required increase significantly
Solution Approach 1:
The patent segments the summarization task into extractive selection of sentences and abstractive generation of only key phrases within those sentences. This hybrid approach extracts relevant sentences using computationally efficient methods, then applies abstractive techniques only to generate concise key phrases, reducing overall computational expense while maintaining readability
Solution Approach 2:
The patent applies different processing qualities to different parts of the summary generation process. Full abstractive summarization is applied locally to generate key phrases within selected sentences, while the overall sentence selection uses efficient extractive methods. This localized application of computationally intensive techniques maintains readability where needed while reducing overall computational burden
2Loss of information
If risk mining is applied to large, high volume data sets, then relevant extractions can be returned, but substantial time is required for an analyst to review the voluminous extractions
Solution Approach 1:
The patent extracts only the most relevant sentences and key phrases from large volumes of data, filtering out unnecessary information. By using entity-risk relationship detection and relevance scoring, the system extracts a concentrated set of high-value extractions that maintain information quality while dramatically reducing the volume requiring analyst review
Solution Approach 2:
The patent performs preliminary automated summarization and ranking of extractions before analyst review. By pre-processing the data to identify and rank the most relevant sentences and key phrases, the system prepares a curated set of high-priority extractions, reducing the time analysts need to spend reviewing voluminous unprocessed data
3Productivity
If a single sentence extraction is used to represent a risk event, then processing is efficient, but the extraction may not provide enough information for an analyst to properly determine relevance
Solution Approach 1:
The patent merges extractive sentence selection with abstractive key phrase generation in a hybrid approach. Multiple relevant sentences are extracted to maintain information completeness, while abstractive key phrases are generated within each sentence to highlight critical information. This combination preserves comprehensive information while maintaining processing efficiency through automated selection and highlighting
Data Source
AI summary
In some aspects, a method includes extracting sentences from data corresponding to documents. Each extracted sentence includes at least one matched pair (a keyword from a first or second keyword set and an entity from an entity set). The method includes ordering the plurality of extracted sentences based on a distance between a respective keyword and a respective entity in each extracted sentence. The method includes identifying a first type and a second type of extracted sentences from the ordered plurality of extracted sentences. Sentences having the first type include keywords of the first keyword set. Sentences having the second type include keywords of the second keyword set. The method includes generating an extracted summary including at least one sentence having the first type and at least one sentence having the second type, intermixed based on a predetermined order rule set. The method includes outputting the extracted summary.


