Mobile Document Image Validation Using On-Device Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing mobile check deposit solutions face challenges with inconsistent image quality due to low light conditions and shaky hands, difficulty in verifying user identity and document authenticity, and heavy reliance on backend servers for processing, leading to delays and resource consumption.

Innovation Solution

A computing system that performs client-side validation using machine-learning models to enhance image quality, detect document features, and verify user identity, reducing the need for server-side processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If mobile devices are used for check deposit, then convenience is improved, but image quality becomes inconsistent

Engineering Contradiction:
ImproveconvenienceVSAvoidimage quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical image capture systems with mobile device cameras with an AI-powered image processing system. The system uses machine learning models to analyze, enhance, and standardize images captured by various mobile devices, compensating for differences in camera quality, lighting conditions, and user handling. This substitution enables consistent image processing despite the diversity of mobile capture devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Power

If backend servers perform image processing, then processing capability is improved, but processing time increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidprocessing time
Core Design Contradiction:
PowerVSLoss of time

Solution Approach 1:

The patent segments the image processing workflow into two parts: initial image analysis and validation are performed locally on the mobile device using embedded AI models, while only essential data and selected images are transmitted to backend servers for final processing. This segmentation reduces the time-critical processing steps to the client side, minimizing latency while maintaining server capabilities for complex tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary image validation, quality assessment, and feature extraction on the mobile device before transmission to the backend. This preliminary action filters out low-quality images and prepares data in advance, reducing the processing burden and time required when images reach the server.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If detailed instructions are provided to users, then verification accuracy is improved, but user convenience decreases

Engineering Contradiction:
Improveverification accuracyVSAvoiduser convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements real-time feedback mechanisms where the AI system analyzes images as they are captured and provides immediate guidance to users through the mobile interface. The system offers contextual, step-by-step instructions based on the specific image quality issues detected, rather than providing generic detailed instructions beforehand. This feedback loop maintains verification accuracy while improving convenience by guiding users only when and where needed.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12586402B2Machine-learning models for image processing
Publication Date: 2026.03.24 CITIBANK N A
  • US12586402B2 patent drawing
  • US12586402B2 patent drawing
  • US12586402B2 patent drawing

AI summary

Presented herein are systems and methods for the employment of machine learning models for image processing as may be performed by computing devices associated with an end user. A method may include obtaining video data comprising a plurality of frames including a document of a document type. The method may include executing an object recognition engine of a machine-learning architecture using image data of the plurality of frames, the object recognition engine trained to detect edges of documents. The method may include identifying, based on the edge detection, a plurality of boundaries for the document. The method may include validating, based on the plurality of boundaries, the document as the document type. The method may include transmitting via one or more networks, to a computer remote from the computing device, responsive to the validation of the type of document, the image data for the plurality of frames depicting the document.