YOLO CNN Document Intake System for Handwritten Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automated document processing systems face challenges in accurately extracting data from handwritten forms due to varying legibility, leading to potential errors in critical fields like healthcare and finance, where precise data integrity is crucial.

Innovation Solution

An automated document intake and processing system utilizing a pre-trained 'YOLO' convolutional neural network (CNN) for image processing, which applies filters to classify and extract target data from forms, incorporating pre-processing steps and machine vision techniques to improve accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If traditional OCR and machine vision processing techniques are used to extract data from handwritten forms, then the intake process can be automated to some extent, but the accuracy and reliability of data extraction deteriorate due to varying legibility of handwriting

Engineering Contradiction:
Improveautomation of data extractionVSAvoiddata extraction accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system performs pre-processing steps on the source image before applying the YOLO CNN model. This includes image enhancement, normalization, and other preparatory actions that improve the quality of input data, thereby enabling more accurate extraction of handwritten information while maintaining automation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional OCR and machine vision algorithms with a deep learning-based YOLO CNN model. This substitution enables the system to learn and adapt to various handwriting styles and legibility levels, significantly improving extraction accuracy while maintaining automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If human intake processors are employed to manually extract and validate data from forms, then data extraction accuracy can be maintained, but processing time and operational costs increase

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The YOLO CNN model is trained to autonomously identify, classify, and extract target data from various document formats without requiring manual intervention. The system self-adjusts to different handwriting styles and document variations, achieving both high accuracy and automated processing speeds

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses a trainable model that can adapt its parameters based on the specific characteristics of the documents being processed. This allows the automated system to maintain high accuracy across different document types and handwriting legibility levels, matching human processor performance while maintaining automation efficiency

Inventive Principle:
Principle #35Parameter changes

3Productivity

If a standardized template approach is used for document processing, then processing efficiency can be improved, but the system's ability to handle customized templates and varied document formats deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidtemplate flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The YOLO CNN model is designed to handle multiple document formats, templates, and data types within a single unified system. It can process standardized templates, customized templates, and even undocumented formats by learning their characteristics, thereby maintaining processing efficiency while achieving high versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs a trainable model that can dynamically adapt to different document formats and templates. Rather than requiring fixed templates, the model learns the structure and characteristics of various document types, enabling efficient processing of both standardized and customized formats

Inventive Principle:
Principle #15Dynamics

4Productivity

If automated processing is implemented without sufficient validation mechanisms, then processing speed increases, but the risk of errors and data integrity issues worsens in high-risk environments

Engineering Contradiction:
Improveprocessing speedVSAvoiddata integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system incorporates validation mechanisms that provide feedback on the extracted data quality. The YOLO CNN model can identify confidence levels for extracted fields, and the system can flag low-confidence extractions for review, thereby maintaining high processing speeds while ensuring data integrity through automated validation

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs validation and verification steps as part of the automated processing pipeline. By incorporating these checks before final data extraction and integration, the system ensures data integrity and accuracy while maintaining efficient automated processing throughout the workflow

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11631266B2Automated document intake and processing system
Publication Date: 2023.04.18 WILCO SOURCE LLC
  • US11631266B2 patent drawing
  • US11631266B2 patent drawing
  • US11631266B2 patent drawing

AI summary

An automated documentation intake and processing system involves pre-processing source images of intake documents and applying a YOLO CNN model to identify fields of interest (FOIs) therein which may contain target data. The system classifies FOIs and upsamples the contents in order to digitize and extract target data. The system may be trained to differentiate between target data and non-target data and incorporates an adjustable confidence scores which reflects the system's degree of accuracy at predicting the correct FOI (e.g., name, address, insurance number, vehicle registration number). The system is pre-trained to detect a subset of documentation types, form fields, and form field types. However, the system is configured to adapt to variations of the same types of documentation, such as different insurance cards or driver licenses from different states.