Document Selection System Optimizing Quality and Diversity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulty in sorting through the large volume of documents from various media sources, particularly non-traditional sources like micro-blogs, due to the overwhelming number of documents produced by multiple authors.

Innovation Solution

A method and system for separating a set of related documents by determining quality scores and similarity scores, and solving an optimization problem to obtain a subset of documents that are both of high quality and diverse, thereby reducing the number of documents while maintaining relevance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If all documents from multiple media sources are collected and presented to users, then the completeness of information is improved, but the complexity of document sorting and user accessibility deteriorates

Engineering Contradiction:
Improvecompleteness of informationVSAvoiddocument sorting accessibility
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the large volume of documents into smaller, manageable subsets based on quality scores and similarity metrics. The system divides documents into different groups (e.g., high-quality unique documents, lower-quality duplicates) and presents them in a structured manner, making it easier for users to navigate and access relevant information without being overwhelmed by the total volume.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and removes duplicate or low-quality documents from the complete set. By calculating similarity scores and quality metrics, the system identifies and extracts only the most valuable unique documents for presentation, while still maintaining access to the full document set if needed. This extraction process reduces the visible document volume while preserving information completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of information

If the complete set of related documents is presented to users, then the comprehensiveness of content is improved, but the time required to process and visualize documents deteriorates

Engineering Contradiction:
Improvecomprehensiveness of contentVSAvoiddocument processing and visualization time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-calculating quality scores and similarity scores for all documents before presentation. The system prepares document subsets in advance based on these pre-computed metrics, so that when users request documents, the system can quickly present pre-processed, organized subsets rather than processing everything in real-time. This reduces visualization time while maintaining comprehensiveness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by presenting a carefully selected subset of documents that captures the essential comprehensiveness without requiring processing of the entire document set. The system calculates that a smaller subset (e.g., top N quality documents with low similarity) provides sufficient comprehensiveness for user needs, avoiding the excessive time required to process and display all documents.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If a large number of documents are displayed to users, then the completeness of information coverage is improved, but the visual clarity and manageability for users deteriorates

Engineering Contradiction:
Improveinformation coverage completenessVSAvoidvisual clarity and manageability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the document display into organized subsets based on quality and similarity criteria. Instead of displaying all documents in a single large list, the system divides them into manageable groups (e.g., high-quality unique documents, related documents with lower similarity) that can be visually processed more easily. This segmentation maintains information coverage while improving visual clarity through structured presentation.

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If all documents from multiple authors are collected, then the diversity of perspectives is improved, but the volume of documents to be processed deteriorates

Engineering Contradiction:
Improvediversity of perspectivesVSAvoiddocument volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts unique high-quality documents that represent diverse perspectives while removing redundant content. By calculating similarity scores between documents, the system identifies and extracts only the most valuable representative documents from multiple authors, preserving perspective diversity while reducing overall document volume. The extraction process ensures that each selected document contributes unique value.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8898151B2System and method for filtering documents
Publication Date: 2014.11.25 ROGERS COMMUNICATIONS
  • US8898151B2 patent drawing
  • US8898151B2 patent drawing
  • US8898151B2 patent drawing

AI summary

A method and document separation system for separating a set of related documents is described. In one aspect, the method comprises: determining, on a document selection system, quality scores for a plurality of the documents in the set of related documents; obtaining a similarity score for a plurality of pairs of documents in the set of related document; and on a document selection system, obtaining a first subset of related documents which solves an optimization problem, the first subset of related documents including a portion of the document in the set of related documents, the optimization problem being a function of one or more quality scores of the documents assigned to the first subset of related documents and one or more similarity scores of pairs of documents assigned to the first subset of related documents.