Unified Vision-Language Model for Document Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document AI processors are limited to specific document types and require extensive manual labeling and training, making them ineffective for layout variations and new document types, and lack the ability to accurately extract information using the two-dimensional layout and appearance of documents.
Innovation Solution
A system that uses a unified language-image model, such as the FormPaLI model, for joint learning of language and vision features, allowing it to process document queries and generate answers with bounding boxes indicating the location of the answer within the document image, improving OCR capabilities and layout understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specialized processors are trained on large amounts of manually labeled data for specific document types, then extraction accuracy for those document types is improved, but the system cannot be applied to different document or entity types without further labeling and training
Solution Approach 1:
The patent applies universality by training a single processor on diverse document types (invoices, receipts, forms, contracts) using manually labeled data, enabling the same processor to handle multiple document types and entity types without requiring separate specialized processors for each document type
Solution Approach 2:
The patent changes the training parameters by using a large dataset spanning multiple document types and entity types with varied layouts and formats, allowing the processor to learn generalizable features rather than document-type-specific patterns
2Productivity
If template-based entity extraction is used for documents following the same layout, then extraction is efficient for consistent layouts, but the template becomes ineffective when layout variations or new layouts are encountered
Solution Approach 1:
The patent applies dynamics by using a machine learning processor that can adapt its extraction behavior based on the input document's layout characteristics, rather than relying on a fixed static template. The processor dynamically adjusts to different layouts through its trained understanding of various document structures
Solution Approach 2:
The processor performs self-service by automatically adapting to new document layouts through its generalizable training, eliminating the need for manual template updates or reconfiguration when encountering new document types or layout variations
3Measurement precision
If custom processors are trained on customer-labeled data for specific use cases, then extraction accuracy for those specific use cases is improved, but the process is time-consuming and expensive
Solution Approach 1:
The patent applies universality by creating a single processor trained on diverse data that can serve multiple customer use cases across different document types and entity types, eliminating the need for customers to train separate custom processors for each use case
Solution Approach 2:
The processor provides self-service capabilities by being pre-trained on comprehensive data that covers various document types and entity types, allowing customers to use it immediately for multiple purposes without investing time and resources in custom training
Data Source
AI summary
An aspect of the disclosed technology is a system and process that are able to answer a document query as text and also provide the location in an image where the answer text is detected. In one aspect of the disclosed technology, a machine learning model combines vision and language features for joint learning.


