Ordered Reading Lists from Unstructured Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual review of clustered documents is inefficient due to redundancy and overlap, requiring readers to spend significant time identifying novel content, often leading to skimming and potential missed information.

Innovation Solution

A clustering process using text analytics and natural language processing to analyze documents, prioritize relevance, and organize them into ordered reading lists, with visual cues to highlight novel information and hide less relevant content, allowing users to customize their reading path based on feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If documents are organized into clustered groups for review, then information organization is improved, but reading efficiency deteriorates due to redundancy and overlap

Engineering Contradiction:
Improveinformation organizationVSAvoidreading efficiency
Core Design Contradiction:
Stability of the object's compositionVSProductivity

Solution Approach 1:

The patent segments documents into clustered groups based on topic similarity, then further segments each document into paragraphs and sentences. This hierarchical segmentation allows the system to identify and highlight only the novel portions (new paragraphs or sentences) within each document, rather than requiring readers to process entire documents. This resolves the contradiction by maintaining organized clustering while eliminating redundant reading through granular segmentation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and highlights only the novel content (new paragraphs or sentences) from each document in the cluster, separating it from redundant information. By extracting only the unique contributions of each document and presenting them prominently, the system maintains the organizational benefits of clustering while dramatically improving reading efficiency by eliminating the need to read through repeated information.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of information

If readers manually review entire documents to determine relevance, then comprehensive understanding is improved, but time consumption deteriorates

Engineering Contradiction:
Improvecomprehensive understandingVSAvoidtime consumption
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis by comparing each document against previously read documents in the cluster to identify novel content before the reader encounters it. New paragraphs and sentences are pre-highlighted, allowing readers to immediately see what information is unique and worth their time. This preliminary identification of novel content maintains comprehensive understanding while dramatically reducing time consumption by eliminating the need to read entire documents.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different visual treatments to different parts of the document based on their novelty status. New paragraphs and sentences are highlighted with distinct visual cues, while redundant content is either suppressed or marked differently. This local differentiation allows readers to focus their attention on high-value information while maintaining the ability to access comprehensive content if needed, thus reducing time consumption without sacrificing understanding.

Inventive Principle:
Principle #3Local quality

3Productivity

If visual cues highlight novel information, then reading efficiency is improved, but system complexity deteriorates

Engineering Contradiction:
Improvereading efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical reading processes with automated text analytics and natural language processing systems. The system automatically compares documents, identifies novel paragraphs and sentences, and applies visual highlighting without human intervention. This substitution of automated computational processes for manual reading analysis improves reading efficiency while the complexity is managed through software rather than requiring complex physical or manual systems.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9454528B2Method and system for creating ordered reading lists from unstructured document sets
Publication Date: 2016.09.27 XEROX CORP
  • US9454528B2 patent drawing
  • US9454528B2 patent drawing
  • US9454528B2 patent drawing

AI summary

A method for creating an ordered reading list for a set of documents includes identifying the topics among documents in a document set; clustering the document set into groups by topic; calculating a probability that a particular topic describes a given document in a cluster based upon the occurrence of the keywords in the document; determining relevant documents in a cluster based on a probability distribution; determining relevant information in a document by repeating a similar operation on the document paragraphs; generating an ordered reading list for the related documents of the cluster based on the relevance; and associating a visual que with non-redundant information in each document to indicate which paragraphs contain the relevant information.