Mobile Document Composite Image Generation via Video Stitching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern mobile devices face challenges in capturing and processing large documents due to hardware limitations, making it difficult to achieve sufficient resolution for downstream processing using traditional image capture and processing algorithms.
Innovation Solution
A system and method for long document stitching using mobile devices, which involves capturing video data, estimating motion vectors, detecting and tracking the document, selecting images based on the tracked position and motion vectors, and generating a composite image to overcome the limitations of capturing large documents in a single image with sufficient resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional image capture algorithms are used on mobile devices, then the device structure remains simple and operation is easy, but the resolution is insufficient for large documents
Solution Approach 1:
The patent segments the capture process into multiple video frames captured sequentially as the mobile device moves over the document. Instead of attempting to capture the entire large document in a single image, the system divides the capture into multiple smaller frame segments that are later stitched together to form a complete high-resolution composite image of the entire document.
2Measurement precision
If multiple images are captured and processed to create composite images, then document resolution improves, but processing time increases
Solution Approach 1:
The patent performs preliminary actions during the capture phase by continuously tracking document position and estimating motion vectors for each video frame as it is captured. This preliminary processing of motion information enables faster subsequent stitching and alignment operations, reducing the overall processing time required to create the composite image.
3Adaptability or versatility
If video data is captured and processed to generate composite images, then document capture capability improves, but computational requirements increase
Solution Approach 1:
The patent extracts and utilizes motion vector information from video frames to guide the composite image generation process. By extracting motion data that describes document position and camera movement, the system can efficiently align and stitch frames without requiring intensive computational processing of entire video sequences, thus reducing computational power requirements while maintaining enhanced capture capability.
Data Source
AI summary
According to one embodiment, a system includes a processor and logic in and/or executable by the processor to cause the processor to: initiate a capture operation using an image capture component of the mobile device, the capture operation comprising; capturing video data; and estimating a plurality of motion vectors corresponding to motion of the image capture component during the capture operation; detect a document depicted in the video data; track a position of the detected document throughout the video data; select a plurality of images using the image capture component of the mobile device, wherein the selection is based at least in part on: the tracked position of the detected document; and the estimated motion vectors; and generate a composite image based on at least some of the selected plurality of images.


