Structured Document Embedding for Accurate Automated Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic document formats, such as DOCX and PDF, store information in an unstructured manner, making it difficult for automated parsing software to accurately interpret and extract relevant data, leading to misread or miscategorized information, which can prevent resumes or documents from reaching human reviewers.
Innovation Solution
A method of creating enhanced documents that encode content in a structured form according to a defined schema, embedding executable code for interaction, and optionally encrypting the content, allowing for accurate parsing and extraction by automated systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Shape
If documents are formatted visually appealing for human readers using traditional formats (DOCX, PDF), then the visual appearance is improved, but the structured information storage is worsened, making automated parsing difficult
Solution Approach 1:
The patent embeds structured data within the visual document format, creating a nested structure where both the visual presentation and structured information coexist. The enhanced document format allows the visual elements to remain appealing while containing embedded structured data that parsers can extract, effectively nesting information within the document without compromising either aspect.
Solution Approach 2:
The invention creates a composite document format that combines traditional visual formatting capabilities with structured data storage. This composite approach allows the document to simultaneously present information visually for human readers while storing it in a structured manner for machine parsing, resolving the contradiction between aesthetic quality and data structure.
2Loss of information
If raw text data formats are used to store structured information according to a defined schema, then machine readability is improved, but the visual presentation capability is worsened
Solution Approach 1:
The enhanced document format is designed to be universal, serving both human readers and machine parsers simultaneously. It maintains the visual presentation capabilities needed for human consumption while incorporating structured data storage that enables automated parsing, making the document format multi-functional rather than serving a single purpose.
Solution Approach 2:
The structured data is nested within the visual document structure, allowing the document to function as both a visual presentation and a structured data container. This nesting enables the same document to be read visually by humans while being parsed by machines, eliminating the need to choose between the two formats.
3Productivity
If automated parsing software extracts information from unstructured documents, then processing speed is improved, but the accuracy of data extraction is worsened due to misread or miscategorized information
Solution Approach 1:
The structured data is prepared in advance during document creation, organizing information according to a defined schema before the document needs to be parsed. This preliminary structuring ensures that when automated parsing occurs, the data is already organized in a way that enables both fast processing and high accuracy, eliminating the need for parsers to struggle with unstructured information.
Solution Approach 2:
The enhanced document format incorporates feedback mechanisms that allow parsing systems to verify and validate extracted information. The structured schema provides reference information that enables the parsing software to check its extractions against expected formats and structures, correcting errors and improving accuracy while maintaining processing speed.
Data Source
AI summary
The disclosed enhanced document creation and parsing systems deal with enhanced documents that allow for the presentation of document content in a preferred visual manner, while ensuring that the document content can be captured accurately by an automated parser with nothing being discarded or misrepresented. The enhanced document creation system may create an enhanced document by encoding document content in accordance with a defined schema, optionally encrypting the resulting structured data into an encrypted byte string, and embedding the encrypted byte string as non-visible metadata in a rendered document. The resulting enhanced document can be completely and accurately parsed by an enhanced document parsing system that is capable of extracting, decrypting and decoding the embedded document metadata.


