Indexing Captioned Objects in Academic Literature

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional abstracting and indexing services fail to effectively index and retrieve data from figures and tables in academic articles, as critical information is often hidden within image files and not searchable, leading to incomplete results in literature searches.

Innovation Solution

A system and method for extracting, indexing, and linking captioned objects such as figures and tables from academic articles, allowing for the capture and retrieval of valuable data through the use of extraction rules, object loaders, and indexing processes, enabling the association of objects with abstracts and full-text records.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional abstracting and indexing services are used, then article-level indexing is provided, but figures and tables containing critical data remain hidden in image files and are not searchable

Engineering Contradiction:
Improvesearchability of figure and table dataVSAvoidindexing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts text and data from figures and tables by processing image files and PDF documents. OCR technology and content extraction algorithms isolate captioned objects and their associated text, separating searchable content from the original image format to make it indexable while maintaining the original document structure.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary indexing layer that bridges traditional article-level indexing and full-text search. This intermediary system processes figures and tables to create searchable metadata and captions, allowing researchers to search for specific data elements without requiring complete full-text indexing of all document content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If full-text indexing is used to index all text within documents, then more text becomes searchable, but key variables are diluted with peripheral matches and precision decreases

Engineering Contradiction:
Improvesearch result precisionVSAvoidcompleteness of searchable content
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies local quality by differentiating between different types of text content within documents. Instead of uniformly indexing all text, the system specifically targets and enhances indexing of figure captions, table titles, and axis labels with higher priority and different indexing strategies than general body text, improving precision for data-related searches.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the document into distinct functional regions: article text, figure captions, table titles, and axis labels. Each segment is processed and indexed differently, with captioned objects receiving specialized attention to extract key variables and data elements, thereby preventing dilution with peripheral text while maintaining completeness.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If web harvesters like Google are used, then all text in documents becomes searchable, but text within images is not extracted and remains invisible to search

Engineering Contradiction:
Improveextractability of image textVSAvoidsearch processing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent performs preliminary extraction of text from images and PDFs during the indexing phase, before search queries are executed. OCR and content extraction algorithms pre-process figure and table images to convert embedded text into searchable metadata, so that when searches occur, the text is already available without requiring real-time image processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7765199B2Method and system to index captioned objects in published literature for information discovery tasks
Publication Date: 2010.07.27 PROQUEST LLC
  • US7765199B2 patent drawing
  • US7765199B2 patent drawing
  • US7765199B2 patent drawing

AI summary

The present invention relates to the identification, extraction, linking, storage and provisioning of data that constitute the captioned components of published or “print ready” literature for computerized information discovery activities including search, browse and data mining. These components, or objects, include the tabular presentation of data (“tables”) and graphics such as “figures”, “images” and “illustrations” typically used to supplement the textual narrative of the publication.