Searchable Annotated Document Data Structures for Citation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic document management systems struggle to analyze documents in their graphic format and generate accurate internal citations, leading to inefficiencies in litigation and patent analysis, where preserving the graphic image while enabling electronic search and citation is crucial.
Innovation Solution
The system converts graphic representations of documents into searchable annotated formats, using data structures to support citation annotation and analysis, allowing for accurate internal citations and thorough term analysis by correlating OCR data with textual versions and generating citation and corpus reports.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If documents are stored as graphic images to preserve original format, then document format preservation is improved, but electronic search capability deteriorates
Solution Approach 1:
The patent creates a text copy (transcription) of the graphic document image. The OCR engine converts the graphic image into searchable text, allowing electronic search without altering the original graphic format. This copy enables full-text search while the original graphic image remains preserved for format stability.
Solution Approach 2:
The patent introduces an intermediary layer (transcribed text with citation data structures) between the graphic image and the search function. This intermediary enables electronic search capability while the graphic image itself remains unchanged, resolving the contradiction between format preservation and searchability.
2Measurement precision
If manual citation creation is used to ensure accuracy, then citation accuracy is improved, but time consumption deteriorates
Solution Approach 1:
The system performs automatic citation generation through OCR transcription and data structure creation. The citation information is extracted and annotated automatically without requiring manual intervention, making the system self-service for citation creation while maintaining accuracy through structured data capture.
Solution Approach 2:
The patent performs preliminary OCR transcription and citation data structure creation during the initial document processing phase. This preliminary action captures all citation information in advance, so that subsequent copying or referencing operations can use pre-validated citation data without requiring additional time for verification.
3Reliability
If comprehensive term analysis is performed to identify all occurrences, then analysis thoroughness is improved, but processing complexity deteriorates
Solution Approach 1:
The patent replaces manual mechanical analysis with automated computational processing. The system uses software to automatically search, identify, and analyze all occurrences of terms throughout the document, providing comprehensive analysis without the complexity of manual review processes.
Solution Approach 2:
The system creates a searchable text copy of the graphic document, enabling automated term analysis. This copy allows for efficient computational searching and analysis of all term occurrences without needing to manually examine the original graphic image, reducing processing complexity while maintaining thoroughness.
Data Source
AI summary
Computer searchable annotated formatted documents are produced by correlating documents stored as a photographic or scanned graphic representations of an actual document (evidence, report, court order, etc.) with textual version of the same documents. A produced document will provide additional details in a computer data structure that supports citation annotation as well as other types of analysis of a document. The computer data structure also supports generation of citation reports and corpus reports. A computer method of creating searchable annotated formatted documents including citation and corpus reports by correlating and correcting text files with photographic or scanned graphic of the original documents. Data structures for correlating and correcting text files with graphic images. Generation of citation reports, concordance reports, and corpus reports. Data structures for citation reports, concordance reports, and corpus reports generation.


