Enhanced Document Metadata for Accurate Automated Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic document formats, such as DOCX and PDF, store information in an unstructured manner, making it difficult for automated parsing software to accurately interpret and extract relevant data, leading to misread or miscategorized information, which can prevent resumes or documents from reaching human reviewers or result in incorrect information being viewed.
Innovation Solution
A method of creating enhanced documents that encode unstructured content in a structured form according to a defined schema and embed non-visible metadata, allowing for accurate parsing by automated systems while maintaining visual appeal for humans.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If documents are stored in visually appealing formats (DOCX, PDF), then human readability is improved, but automated parsing accuracy deteriorates
Solution Approach 1:
The patent embeds structured data within the visual document format, creating a nested structure where the visual representation contains both its own data and an embedded structured data layer. This allows the document to serve dual purposes: visual consumption and machine parsing, resolving the contradiction between human readability and parsing accuracy.
Solution Approach 2:
The patent introduces an intermediary structured data layer between the visual document format and the parsing system. This intermediary layer translates the visual document into machine-readable structured data, enabling accurate parsing while preserving the original visual appeal for human readers.
2Measurement precision
If information is stored in structured formats (XML, JSON), then parsing accuracy is improved, but visual presentation capability deteriorates
Solution Approach 1:
The patent creates a universal document format that serves multiple functions simultaneously: visual presentation for humans and structured data storage for machines. The enhanced document format can be both displayed to users and parsed by automated systems, eliminating the need to choose between visual quality and parsing accuracy.
Solution Approach 2:
The structured data is nested within the visual document structure, allowing the same document to provide both visual presentation and machine-readable data. The visual elements contain embedded structured data that can be extracted without affecting the visual appearance.
3Ease of operation
If resumes are formatted precisely for visual appeal, then human review quality is improved, but automated filtering reliability deteriorates
Solution Approach 1:
The structured data embedded in the resume acts as an intermediary that bridges human review and automated filtering. The visual resume maintains its appeal for human readers while the embedded structured data provides reliable information for automated ATS filtering, ensuring both functions work together reliably.
Data Source
AI summary
The disclosed enhanced document creation and parsing systems deal with enhanced documents that allow for the presentation of document content in a preferred visual manner, while ensuring that the document content can be captured accurately by an automated parser with nothing being discarded or misrepresented. The enhanced document creation system may create an enhanced document by encoding document content in accordance with a defined schema, optionally encrypting the resulting structured data into an encrypted byte string, and embedding the encrypted byte string as non-visible metadata in a rendered document. The resulting enhanced document can be completely and accurately parsed by an enhanced document parsing system that is capable of extracting, decrypting and decoding the embedded document metadata.


