Automated Document Reference Extraction System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Accessing external content referenced in documents is cumbersome, requiring users to manually follow links or search engines, making retrieval tedious.

Innovation Solution

Automated document analysis systems that exploit the inherent layout of structured documents to extract and retrieve referenced external content by correcting artifacts, binarizing images, and analyzing layout to identify regions of interest containing references, allowing for automatic extraction and retrieval of linked content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual processes are used to access external content, then users can control the retrieval process, but the process becomes tedious and time-consuming

Engineering Contradiction:
Improveease of retrievalVSAvoidretrieval time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system automatically extracts references from document images and retrieves external content without requiring user intervention. The automated reference extraction system processes document images, identifies references, and fetches external content autonomously, eliminating the need for manual link following or search engine queries.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (typing URLs, clicking links, using search engines) with an automated computational system that uses image processing, reference detection algorithms, and automated web scraping to retrieve external content, thereby eliminating tedious manual operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated extraction is implemented, then retrieval efficiency improves, but system complexity increases

Engineering Contradiction:
Improveretrieval efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the complex task of external content retrieval into separate modular components: document image processing, reference extraction, content retrieval, and presentation. This segmentation allows each component to be developed and optimized independently, managing overall system complexity while maintaining high productivity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8805074B2Methods and systems for automatic extraction and retrieval of auxiliary document content
Publication Date: 2014.08.12 SHARP KK
  • US8805074B2 patent drawing
  • US8805074B2 patent drawing
  • US8805074B2 patent drawing

AI summary

Aspects of the present invention are related to systems and methods for automatically extracting, from a document image, references to relevant external content and automatically retrieving the external content associated with the references.