User-Driven Document Data Collection via OCR Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional forms-based data collection systems, such as tax return preparation software, often ask irrelevant questions, confuse users with complex terminology, and require documents to be entered in a non-intuitive order, leading to user frustration and potential errors.

Innovation Solution

A user-driven document-based data collection system that allows users to enter data from documents in any order, identifies relevant documents through descriptions or scanned images, and uses OCR to analyze and determine necessary information for specific tasks, such as tax return preparation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional forms-based data collection systems ask every possible question to ensure complete data collection, then data completeness is improved, but user confusion and time consumption increase

Engineering Contradiction:
Improvedata completenessVSAvoiduser time consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically analyzing uploaded documents to extract data before presenting questions to the user. The document analysis component processes documents in advance to identify relevant information, reducing the number of questions the user needs to answer manually while ensuring data completeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service by allowing users to upload documents and having the system automatically extract and populate data fields. Users no longer need to manually answer numerous questions; instead, the system serves itself by extracting information from documents and presenting only relevant questions that require user input.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If traditional forms-based systems use specific industry terminology to ensure accuracy, then data precision is improved, but user accessibility deteriorates

Engineering Contradiction:
Improvedata accuracyVSAvoiduser accessibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system introduces an intermediary layer in the form of a document analysis component that translates complex industry terminology into user-friendly interfaces. This intermediary automatically processes documents and presents extracted information in a simplified manner, maintaining data accuracy while improving user accessibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a copy of the data extraction process that operates independently of traditional question formats. Instead of asking users to interpret and answer terminology-heavy questions, the system copies the data extraction function to automatically process documents, preserving accuracy while eliminating the barrier of complex terminology.

Inventive Principle:
Principle #26Copying

3Stability of the object's composition

If traditional systems require documents to be entered in a specific order to match government forms, then data organization is improved, but user friendliness deteriorates

Engineering Contradiction:
Improvedata organizationVSAvoiduser friendliness
Core Design Contradiction:
Stability of the object's compositionVSEase of operation

Solution Approach 1:

The system introduces dynamics by allowing users to upload documents in any order rather than being constrained by a fixed sequence. The document analysis component dynamically processes documents regardless of upload order, and the system automatically organizes extracted data into the correct structure, maintaining data organization while significantly improving user friendliness.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system inverts the traditional approach by instead of requiring users to follow a predetermined order, the system adapts to the user's natural upload sequence. The data organization function is inverted from a rigid template-based approach to a flexible, user-driven process where the system automatically structures data regardless of input order.

Inventive Principle:
Principle #13The other way round (Inversion)

4Adaptability or versatility

If traditional forms ask repetitive Yes/No questions to cover all possibilities, then data coverage is improved, but user frustration increases

Engineering Contradiction:
Improvedata coverageVSAvoiduser frustration
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system extracts the repetitive question-asking function and replaces it with automated document analysis. By taking out the manual question-response interaction and extracting data directly from documents, the system maintains comprehensive data coverage while eliminating the frustration of repetitive Yes/No questions.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system substitutes the mechanical system of repetitive manual questions with an automated document processing mechanism. Instead of mechanically asking users the same questions repeatedly, the system uses optical character recognition and document analysis to automatically extract data, maintaining coverage while reducing user frustration.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach reduces user confusion by allowing flexible data entry, identifies relevant documents accurately, and streamlines the data collection process, ensuring all necessary information is captured efficiently.

Implementation Method 1

the data collection system may include optical character recognition (OCR) software and may perform various OCR functions on the scanned image to identify the physical document

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS7930226B1User-driven document-based data collection
Publication Date: 2011.04.19 INTUIT INC
  • US7930226B1 patent drawing
  • US7930226B1 patent drawing
  • US7930226B1 patent drawing

AI summary

A user-driven document data collection system may allow the user to enter document data in no particular order. The data collection system may help the user identify documents and determine whether those documents are relevant. The document data collection system may also allow the user to input a description of a document, identify the document based on the description, and determine whether or not the document is appropriate for data collection. The data collection system may be configured to display example documents for the user to verify the identification of a document. A user may enter data for a document via a data entry screen based in part on a scanned image of the document. The document data collection system may analyze the data from documents to determine whether or not any additional information, such as from additional documents, is required to perform a particular task using the document data.