Unified Vision-Language Model for Document Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document AI processors are limited to specific document types and require extensive manual labeling and training, making them ineffective for layout variations and new document types, and lack the ability to accurately extract information using the two-dimensional layout and appearance of documents.

Innovation Solution

A system that uses a unified language-image model, such as the FormPaLI model, for joint learning of language and vision features, allowing it to process document queries and generate answers with bounding boxes indicating the location of the answer within the document image, improving OCR capabilities and layout understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specialized processors are trained on large amounts of manually labeled data for specific document types, then extraction accuracy for those document types is improved, but the system cannot be applied to different document or entity types without further labeling and training

Engineering Contradiction:
Improveextraction accuracyVSAvoidapplicability to different document types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by training a single processor on diverse document types (invoices, receipts, forms, contracts) using manually labeled data, enabling the same processor to handle multiple document types and entity types without requiring separate specialized processors for each document type

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the training parameters by using a large dataset spanning multiple document types and entity types with varied layouts and formats, allowing the processor to learn generalizable features rather than document-type-specific patterns

Inventive Principle:
Principle #35Parameter changes

2Productivity

If template-based entity extraction is used for documents following the same layout, then extraction is efficient for consistent layouts, but the template becomes ineffective when layout variations or new layouts are encountered

Engineering Contradiction:
Improveextraction efficiencyVSAvoidhandling of layout variations
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by using a machine learning processor that can adapt its extraction behavior based on the input document's layout characteristics, rather than relying on a fixed static template. The processor dynamically adjusts to different layouts through its trained understanding of various document structures

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The processor performs self-service by automatically adapting to new document layouts through its generalizable training, eliminating the need for manual template updates or reconfiguration when encountering new document types or layout variations

Inventive Principle:
Principle #25Self-service

3Measurement precision

If custom processors are trained on customer-labeled data for specific use cases, then extraction accuracy for those specific use cases is improved, but the process is time-consuming and expensive

Engineering Contradiction:
Improveuse case accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies universality by creating a single processor trained on diverse data that can serve multiple customer use cases across different document types and entity types, eliminating the need for customers to train separate custom processors for each use case

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The processor provides self-service capabilities by being pre-trained on comprehensive data that covers various document types and entity types, allowing customers to use it immediately for multiple purposes without investing time and resources in custom training

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240394284A1Query-Based Document Extraction with Large Vision-Language Models
Publication Date: 2024.11.28 GOOGLE LLC
  • US20240394284A1 patent drawing
  • US20240394284A1 patent drawing
  • US20240394284A1 patent drawing

AI summary

An aspect of the disclosed technology is a system and process that are able to answer a document query as text and also provide the location in an image where the answer text is detected. In one aspect of the disclosed technology, a machine learning model combines vision and language features for joint learning.