Document Data Extraction Template Ranking and Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems face challenges in providing users with functionality and features without significant user data entry or interaction, particularly due to complexities in data extraction from various document types, which often require custom data extraction templates and lack efficient management mechanisms for these templates.

Innovation Solution

A process for document data extraction template management that calculates a ranking score based on user acceptance and field hit count to prioritize and manage data extraction templates, allowing for dynamic ranking and storage of templates associated with specific source document types, thereby improving efficiency and relevance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data extraction templates are created for every type and format of document, then data extraction accuracy is improved, but device complexity and storage requirements increase significantly

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidtemplate management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal template management system that can handle multiple document types and formats through a single standardized interface. The template data structure is designed to be format-agnostic, allowing the same template framework to extract data from invoices, receipts, tax documents, and other financial documents without requiring separate specialized templates for each format.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts template parameters based on document characteristics. Instead of creating fixed templates for every document type, the system modifies template parameters (such as field locations, data patterns, and extraction rules) to adapt to different document formats, reducing the need for extensive template proliferation.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If users contribute to creating data extraction templates, then template coverage for unknown document formats is improved, but template quality and completeness may deteriorate due to voluntary participation

Engineering Contradiction:
Improvetemplate coverageVSAvoidtemplate quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where extracted data is validated against expected data patterns and formats. User-contributed templates are automatically tested with sample documents, and feedback is provided on template effectiveness. This allows the system to identify and correct quality issues in user-contributed templates while maintaining the benefits of broad user participation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system merges multiple user-contributed templates for the same document type into a consolidated, optimized template. By combining contributions from multiple users and applying automated validation, the system creates higher-quality templates that leverage collective user knowledge while eliminating individual errors and inconsistencies.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If multiple data extraction templates are maintained for the same document type, then extraction accuracy for varying document formats is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements dynamic template selection and generation. Instead of statically maintaining multiple templates, the system dynamically creates or selects the most appropriate template based on the actual document being processed. This dynamic approach reduces the number of templates that need to be evaluated during processing while maintaining high extraction accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary analysis of the document to identify its type and characteristics before full data extraction begins. This preliminary action allows the system to pre-select or pre-generate the appropriate template, avoiding the need to process multiple templates during the main extraction phase and thereby reducing overall processing time.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If comprehensive data extraction templates are created to extract all required fields, then data completeness is improved, but user data entry burden increases during template creation

Engineering Contradiction:
Improvedata completenessVSAvoiduser effort in template creation
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent implements self-service template creation capabilities where the system automatically generates draft templates by analyzing sample documents. Users can then review and refine these automatically generated templates rather than creating them from scratch. This self-service approach significantly reduces user effort while maintaining comprehensive data extraction coverage.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary template generation by automatically analyzing document structures and extracting common data patterns before presenting the template to the user for review. This preliminary action completes much of the template creation work automatically, reducing the manual effort required while ensuring all necessary data fields are captured.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9292579B2Method and system for document data extraction template management
Publication Date: 2016.03.22 INTUIT INC
  • US9292579B2 patent drawing
  • US9292579B2 patent drawing
  • US9292579B2 patent drawing

AI summary

User acceptance of a given data extraction template and the number of data fields that the data extraction template can extract accurately is used to calculate data extraction template ranking, or a ranking score, to be associated with the data extraction template. Then the data extraction template having the highest data extraction template ranking score is used in a first attempt to extract data from a source documents of the source document type associated with the data extraction templates. As more data extraction templates associated with a given source document type are received, data extraction template ranking scores are updated/modified, and, in one example, the data extraction templates having the lowest data extraction template ranking scores are detected/eliminated.