Non-text Object Metadata for Electronic Document Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic document management systems face challenges in locating original documents when users only have a physical or non-native copy, as simple text searches are insufficient for documents lacking text or with non-text objects, leading to inefficiencies and inaccuracies in document retrieval.
Innovation Solution
A method and system that generates metadata for non-text objects from a physical document scan, allowing for a text-based search across a data repository to locate the original electronic document by comparing the metadata with searchable metadata associated with electronic documents, ensuring accurate document retrieval even when the original document lacks text or contains common words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If simple text search is used to locate documents in a data repository, then the search process is simple and fast, but the search fails for documents lacking text or with non-text objects
Solution Approach 1:
The patent segments the document into text components and non-text object components. Each component is processed separately with appropriate search methods - text is searched using OCR and text search, while non-text objects are searched using image recognition and metadata comparison. This segmentation allows the system to handle diverse document types effectively.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the physical document and the search system. Non-text objects are converted into structured metadata (descriptive information, composition information, hierarchical information) that can be searched and compared systematically. This intermediary enables reliable searching of non-text content without requiring complex image processing during search operations.
2Ease of operation
If OCR software is used to convert hardcopy text for searching, then text-based search becomes possible, but the search returns too many results when the document contains only common words
Solution Approach 1:
The patent merges multiple search criteria - text content from OCR, descriptive metadata from non-text objects, and hierarchical metadata from document structure. By combining these different types of information, the system creates a multi-dimensional search approach that significantly reduces false positives and improves result precision compared to text-only search.
Solution Approach 2:
The patent adds new dimensions to the search space by incorporating metadata dimensions (descriptive, compositional, hierarchical) beyond the traditional text dimension. This multi-dimensional approach allows the system to distinguish between documents that may share common text but differ in their non-text content or structure, thereby improving search precision.
3Productivity
If a user recreates a document when unable to locate the original, then the user can proceed with work, but the reconstructed document may not be identical to the original electronic document
Solution Approach 1:
The patent implements a feedback mechanism where the system compares the physical document's metadata against the metadata of documents in the data repository. This comparison provides feedback to the user about the likelihood of finding the original, allowing the user to decide whether to continue searching or proceed with recreation, thereby optimizing the balance between productivity and document fidelity.
Solution Approach 2:
The patent performs preliminary actions by pre-processing documents in the data repository - extracting metadata from non-text objects and organizing it into searchable structures before the search is needed. This preliminary preparation enables rapid and accurate comparison when a search is initiated, reducing the time users spend searching and the need for document recreation.
4Reliability
If metadata is generated and compared for non-text objects, then document discovery accuracy is improved, but the system complexity increases
Solution Approach 1:
The patent segments metadata generation into distinct functional modules: descriptive metadata extraction (what the object is), compositional metadata extraction (how the object is constructed), and hierarchical metadata extraction (where the object is positioned). Each module handles a specific aspect of metadata generation, making the overall complex process more manageable and maintainable.
Solution Approach 2:
The patent creates a universal metadata framework that handles multiple types of non-text objects (images, charts, graphs, diagrams) using a common structure and approach. The metadata schema is designed to be adaptable to different object types while maintaining consistency, allowing the system to handle diverse content types without requiring separate complex processing paths for each type.
Data Source
AI summary
A method for document discovery includes receiving a scan of a physical copy of a document with a non-text object, determining a tag for the non-text object defining a portion of the non-text object in an original file, and generating, based on the tag, non-text object metadata with composition information of the non-text object. The method further includes searching, using the non-text object metadata, electronic documents stored in a data repository, where each of the electronic documents has an object and searchable metadata associated with the object, comparing the non-text object metadata with the searchable metadata, and providing a location of the original file to a user when the non-text object metadata matches the searchable metadata.


