Document Summarization via Iterative Vector Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information processing systems face challenges in efficiently managing and summarizing large volumes of unstructured text data, as manual screening and rule-based methods are tedious and time-consuming, limiting the scalability of data analysis.

Innovation Solution

An apparatus and method for document summarization through iterative filtering of unstructured text data, using vector representations to determine similarity and iteratively add relevant portions to the summary, with designated stopping criteria to generate a final summary efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual screening methods are used to process unstructured text data, then processing accuracy can be maintained, but processing time and resource consumption increase significantly

Engineering Contradiction:
Improveprocessing accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments unstructured text data into structured formats (e.g., parsing documents into fields, entities, and relationships), enabling automated processing while maintaining accuracy. This segmentation allows the system to handle large volumes of data efficiently without requiring manual review of each document.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that includes validation rules, transformation logic, and quality assurance mechanisms. This intermediary layer automatically verifies and cleanses data during the conversion from unstructured to structured format, maintaining high accuracy without manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Extent of automation

If rule-based methods are used for unstructured text data processing, then data analysis can be automated, but system complexity and maintenance burden increase

Engineering Contradiction:
Improveautomation levelVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent implements a universal structured data processing framework that can handle multiple types of unstructured data (documents, emails, reports) through a single standardized interface. This multi-functional approach automates processing across different data types without requiring separate complex rule sets for each format.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms unstructured text data into structured formats with defined parameters and schemas. By changing the data representation from free-text to parameterized structures, the system achieves automation through standardized data models that are easier to process and maintain than complex text-based rule systems.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If conventional summarization methods are used, then critical information can be identified, but processing speed decreases for large volumes of data

Engineering Contradiction:
Improveinformation retentionVSAvoidprocessing speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent performs preliminary structuring and validation of text data before summarization, organizing content into standardized formats with identified key fields and entities. This preliminary organization enables faster extraction of critical information during summarization, as the data is already arranged in an accessible structure rather than requiring analysis of raw unstructured text.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates structured copies of unstructured data with preserved critical information in standardized formats. These structured representations maintain the essential content and relationships while enabling rapid processing and summarization, achieving both information retention and processing speed.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230116115A1Document summarization through iterative filtering of unstructured text data of documents
Publication Date: 2023.04.13 DELL PROD LP
  • US20230116115A1 patent drawing
  • US20230116115A1 patent drawing
  • US20230116115A1 patent drawing

AI summary

An apparatus comprises at least one processing device configured to receive a query to generate a summary of a document, to perform two or more iterations of filtering the document to produce a current version of the summary of the document, wherein each of the iterations comprises determining similarity between a first vector representation of the current version of the summary and second vector representations of respective ones of two or more portions of the unstructured text data of the document not yet added to the current version of the summary. The processing device is also configured to generate, following identification of one or more designated stopping criteria in a given iteration, a final version of the summary based at least in part on the current version of the summary produced in the given iteration, and to provide a response to the query comprising the final version of the summary.