Document Codification for Accurate Information Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in automatically processing electronic documents, such as invoices, which vary widely in form and content, making it difficult to accurately identify and extract relevant information due to the lack of standardization in presentation and information content.
Innovation Solution
The system generates a numerical codification of document features, referred to as a document codification, which assists in identifying similar documents and extracting information by comparing attributes across a reference set of codifications, allowing for the identification and extraction of canonical features like dates, entities, and amounts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional document processing methods are used, then information can be identified from documents, but processing accuracy deteriorates due to wide variation in document forms and content presentation
Solution Approach 1:
The patent transforms document features into numerical parameters through codification. Document attributes such as text content, position, font size, and layout are converted into numerical values that can be processed and compared systematically, enabling accurate information identification despite format variations.
Solution Approach 2:
The patent creates a universal codification framework that can handle multiple document types and formats. By establishing a standardized numerical representation system, the solution achieves versatility across different document forms while maintaining consistent processing accuracy.
2Adaptability or versatility
If detailed document analysis is performed to handle format variations, then adaptability improves, but processing time and computational complexity increase
Solution Approach 1:
The patent performs preliminary codification of document features into numerical parameters before actual information extraction. This pre-processing step organizes document characteristics into a standardized format, making subsequent analysis faster and more efficient despite the need to handle format variations.
Solution Approach 2:
The patent divides document analysis into separate codification components, each handling specific document features independently. This segmentation allows parallel processing of different document attributes, reducing overall processing time while maintaining comprehensive format adaptability.
3Productivity
If numerical codification is used to improve processing efficiency, then processing speed improves, but information loss may occur during abstraction
Solution Approach 1:
The patent uses numerical parameters to represent document features without losing essential information. Each codification parameter captures specific document characteristics (position, size, content) in a compressed numerical form that retains all necessary details for accurate information extraction while enabling fast processing.
Data Source
AI summary
A computer-implemented method comprises defining a set of canonical features for a document type and a plurality of attributes for a canonical feature; identifying a set of text rectangles from an electronic document; obtaining a comparison set of reference document codifications, one of which comprising a plurality of canonical feature codifications, one of which comprising one or more attribute values for one or more of the plurality of attributes of one of the set of canonical features as the one canonical feature appears in the one reference document; for each current canonical feature of the set of canonical features: selecting a set of canonical feature codifications from the comparison set and identifying a match between one of the set of text rectangles and one of the set of canonical feature codifications; for each of the set of text rectangles, selecting one of the matching canonical feature codifications.


