Indexing Captioned Objects in Academic Literature
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional abstracting and indexing services fail to effectively index and retrieve data from figures and tables in academic articles, as critical information is often hidden within image files and not searchable, leading to incomplete results in literature searches.
Innovation Solution
A system and method for extracting, indexing, and linking captioned objects such as figures and tables from academic articles, allowing for the capture and retrieval of valuable data through the use of extraction rules, object loaders, and indexing processes, enabling the association of objects with abstracts and full-text records.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional abstracting and indexing services are used, then article-level indexing is provided, but figures and tables containing critical data remain hidden in image files and are not searchable
Solution Approach 1:
The patent extracts text and data from figures and tables by processing image files and PDF documents. OCR technology and content extraction algorithms isolate captioned objects and their associated text, separating searchable content from the original image format to make it indexable while maintaining the original document structure.
Solution Approach 2:
The patent introduces an intermediary indexing layer that bridges traditional article-level indexing and full-text search. This intermediary system processes figures and tables to create searchable metadata and captions, allowing researchers to search for specific data elements without requiring complete full-text indexing of all document content.
2Measurement precision
If full-text indexing is used to index all text within documents, then more text becomes searchable, but key variables are diluted with peripheral matches and precision decreases
Solution Approach 1:
The patent applies local quality by differentiating between different types of text content within documents. Instead of uniformly indexing all text, the system specifically targets and enhances indexing of figure captions, table titles, and axis labels with higher priority and different indexing strategies than general body text, improving precision for data-related searches.
Solution Approach 2:
The patent segments the document into distinct functional regions: article text, figure captions, table titles, and axis labels. Each segment is processed and indexed differently, with captioned objects receiving specialized attention to extract key variables and data elements, thereby preventing dilution with peripheral text while maintaining completeness.
3Loss of information
If web harvesters like Google are used, then all text in documents becomes searchable, but text within images is not extracted and remains invisible to search
Solution Approach 1:
The patent performs preliminary extraction of text from images and PDFs during the indexing phase, before search queries are executed. OCR and content extraction algorithms pre-process figure and table images to convert embedded text into searchable metadata, so that when searches occur, the text is already available without requiring real-time image processing.
Data Source
AI summary
The present invention relates to the identification, extraction, linking, storage and provisioning of data that constitute the captioned components of published or “print ready” literature for computerized information discovery activities including search, browse and data mining. These components, or objects, include the tabular presentation of data (“tables”) and graphics such as “figures”, “images” and “illustrations” typically used to supplement the textual narrative of the publication.


