Digital Text Corpus Concept Browsing via Key Term Indexing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Digital text corpora lack a link structure, making it difficult to browse documents by related concepts or characteristics, as opposed to traditional web browsing, due to the rarity of functional hypertext references and citations.

Innovation Solution

A computer-implemented method to identify key terms related to similar passages within a digital text corpus by analyzing contexts across multiple documents, using a passage mining engine to identify similar passages and a key term generation engine to extract and store key terms, enabling navigation and retrieval of related information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If digital text corpus is stored without link structure, then storage simplicity is maintained, but browsing capability by related concepts deteriorates

Engineering Contradiction:
Improvestorage simplicityVSAvoidbrowsing capability
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The patent segments the browsing function from the storage structure by introducing a separate indexing system. The corpus remains stored as simple documents, while a parallel index structure extracts and organizes key terms and passages independently, enabling efficient browsing without complicating the underlying storage architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary indexing system that mediates between the simple stored corpus and the user's browsing needs. The index acts as a mediator that creates logical connections between unrelated documents through extracted key terms and similar passages, enabling concept-based browsing without adding physical links to the stored documents.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If functional hypertext references are added to enable browsing by related concepts, then browsing capability is improved, but document complexity and processing difficulty increase

Engineering Contradiction:
Improvebrowsing capabilityVSAvoiddocument complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent separates the browsing infrastructure from the document content by creating an independent index structure. Instead of embedding references within documents, the system extracts key terms and creates a separate passage index that maps relationships between documents, thereby improving browsing capability without increasing document complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary indexing system that provides browsing functionality without modifying the original documents. The index serves as a mediator layer that creates logical connections between documents based on extracted key terms and similar passages, enabling concept-based navigation while keeping the document structure simple and unmodified.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If citations and inline references are added to support link-based browsing, then browsing capability is improved, but mining difficulty and processing time increase

Engineering Contradiction:
Improvebrowsing capabilityVSAvoidmining time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-extracting key terms and similar passages from documents during an indexing phase. This advance processing creates a ready-to-use index structure that enables rapid browsing without requiring real-time analysis of document content, thereby improving browsing capability while minimizing the time cost during actual use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts essential browsing information (key terms and similar passages) from the full document content during indexing. By taking out only the necessary elements for browsing and storing them in an optimized index structure, the system enables fast concept-based navigation without requiring processing of the complete document texts during browsing operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9323827B2Identifying key terms related to similar passages
Publication Date: 2016.04.26 GOOGLE LLC
  • US9323827B2 patent drawing
  • US9323827B2 patent drawing
  • US9323827B2 patent drawing

AI summary

Key terms for similar passages from a large corpus are identified and used to enhance searching and browsing the corpus. The corpus contains multiple documents such as the text of books. Browsing by concept is supported by identifying a set of similar passages or quotations in documents stored in the corpus and assigning key terms to passages which links conceptually related passages together. The context of each passage instance is identified and can include, for example, the text surrounding the passage. The contexts of all similar passage instances are analyzed in order to identify key terms for the similar passage. The related key terms are analyzed to identify relationships among the key terms from different similar passage sets. The key terms can be used as a basis for navigating the documents in the corpus. The key terms enable browsing the documents in the corpus by concepts referenced in the documents.