Mobile Document Composite Image Generation via Video Stitching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern mobile devices face challenges in capturing and processing large documents due to hardware limitations, making it difficult to achieve sufficient resolution for downstream processing using traditional image capture and processing algorithms.

Innovation Solution

A system and method for long document stitching using mobile devices, which involves capturing video data, estimating motion vectors, detecting and tracking the document, selecting images based on the tracked position and motion vectors, and generating a composite image to overcome the limitations of capturing large documents in a single image with sufficient resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional image capture algorithms are used on mobile devices, then the device structure remains simple and operation is easy, but the resolution is insufficient for large documents

Engineering Contradiction:
Improvedocument resolutionVSAvoidcapture system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the capture process into multiple video frames captured sequentially as the mobile device moves over the document. Instead of attempting to capture the entire large document in a single image, the system divides the capture into multiple smaller frame segments that are later stitched together to form a complete high-resolution composite image of the entire document.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple images are captured and processed to create composite images, then document resolution improves, but processing time increases

Engineering Contradiction:
Improvedocument resolutionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions during the capture phase by continuously tracking document position and estimating motion vectors for each video frame as it is captured. This preliminary processing of motion information enables faster subsequent stitching and alignment operations, reducing the overall processing time required to create the composite image.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If video data is captured and processed to generate composite images, then document capture capability improves, but computational requirements increase

Engineering Contradiction:
Improvedocument capture capabilityVSAvoidcomputational power
Core Design Contradiction:
Adaptability or versatilityVSPower

Solution Approach 1:

The patent extracts and utilizes motion vector information from video frames to guide the composite image generation process. By extracting motion data that describes document position and camera movement, the system can efficiently align and stitch frames without requiring intensive computational processing of entire video sequences, thus reducing computational power requirements while maintaining enhanced capture capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10108860B2Systems and methods for generating composite images of long documents using mobile video data
Publication Date: 2018.10.23 TUNGSTEN AUTOMATION CORPORATION
  • US10108860B2 patent drawing
  • US10108860B2 patent drawing
  • US10108860B2 patent drawing

AI summary

According to one embodiment, a system includes a processor and logic in and/or executable by the processor to cause the processor to: initiate a capture operation using an image capture component of the mobile device, the capture operation comprising; capturing video data; and estimating a plurality of motion vectors corresponding to motion of the image capture component during the capture operation; detect a document depicted in the video data; track a position of the detected document throughout the video data; select a plurality of images using the image capture component of the mobile device, wherein the selection is based at least in part on: the tracked position of the detected document; and the estimated motion vectors; and generate a composite image based on at least some of the selected plurality of images.