Distillate Engine for Interactive Text Navigation and Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Searching and retrieving relevant information from large digital content, such as PDF documents, is laborious and time-consuming due to the need to sift through extensive material, even with search engines, which present raw documents for further analysis.

Innovation Solution

A system that processes text from selected document collections, performs analysis including punctuation, word frequency, and phrase detection, and presents a distilled data set through APIs, allowing users to navigate and extract information interactively without opening documents, using a submission interface, distillate engine, and distillate browser for intuitive access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If search engines present raw documents for retrieval, then information completeness is improved, but information accessibility deteriorates

Engineering Contradiction:
Improveinformation completenessVSAvoidinformation accessibility
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system extracts key information elements (entities, relationships, concepts) from raw documents and presents them in a structured, navigable format. The distillate engine processes full documents to extract essential information without requiring users to view complete raw documents, thus maintaining information completeness while improving accessibility.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing layer (distillate engine) between the raw document repository and the user interface. This intermediary transforms unstructured documents into structured distillates with entities, relationships, and concepts, enabling users to access information through multiple navigation paths without directly handling raw documents.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If users manually search through large documents, then information accuracy is improved, but time consumption increases

Engineering Contradiction:
Improveinformation accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of documents before user queries, pre-extracting entities, relationships, and concepts into a structured distillate format. When users search, they query the pre-processed distillate rather than scanning raw documents, maintaining accuracy through structured data while dramatically reducing search time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical searching through documents with automated computational processing. The distillate engine automatically processes documents and enables semantic navigation through programmed algorithms, substituting human effort with automated systems that maintain accuracy while reducing time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If the system processes and structures text from multiple documents, then information accessibility is improved, but system complexity increases

Engineering Contradiction:
Improveinformation accessibilityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system segments the complex task of document processing into distinct functional modules: text extraction, entity recognition, relationship identification, concept mapping, and result presentation. Each module handles a specific aspect of processing, making the overall complex system manageable and maintainable while delivering improved information accessibility.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8280878B2Method and apparatus for real time text analysis and text navigation
Publication Date: 2012.10.02 EBRARY
  • US8280878B2 patent drawing
  • US8280878B2 patent drawing
  • US8280878B2 patent drawing

AI summary

An end user, by way of a submission interface, instructs an engine to select particular collections of documents to process. The engine processes all the text from within all the documents from within the selected collections. The result of the processing of such text is a distilled data set. Such distillate data set is accessed through APIs by a browser. Different views of the accessed distillate data set may be presented to the end user via the browser allowing them to more effectively assess the utility of the presented data and thereby responsively tune the presented data set with regard to their particular research task. One or more of such views may be used to create a new document from sentences, paragraphs, chapters or documents from the distillate data set that correspond to the one or more views for presentation to the end user.