Document Mosaicing via Video Stream Homography
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for capturing and translating large documents require multiple still images and manual alignment, leading to potential blurring and inefficiencies in processing.
Innovation Solution
A method and system for reconstructing a document mosaic from video streams by identifying feature points, calculating homography mappings between video frames, and rendering an assembled image that depicts the full view of a document, using techniques such as feature point descriptors and binning to improve image alignment and processing speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If multiple still images are captured to cover a large document, then the full document view is obtained, but the processing time and complexity increase
Solution Approach 1:
The patent segments the document capture process into multiple video frames captured in rapid succession, then processes these frames through feature point detection and homography calculation to assemble a complete document view. This segmentation approach allows parallel processing of multiple frames simultaneously, reducing overall processing time compared to sequential still image capture and alignment.
Solution Approach 2:
The patent performs preliminary actions by capturing multiple video frames with inherent temporal overlap before processing begins. The video capture phase pre-establishes the spatial relationships between document regions, and the subsequent processing uses pre-computed feature point descriptors and homography matrices to rapidly assemble the complete document view, avoiding time-consuming real-time alignment operations.
2Area of stationary object
If multiple still images are captured manually aligned, then the full document view is obtained, but image blurring occurs due to device movement
Solution Approach 1:
The patent replaces manual mechanical alignment operations with automated computational methods. Instead of requiring users to manually position and align still images, the system uses computer vision algorithms to detect feature points, compute homography transformations, and automatically stitch multiple video frames together, eliminating blurring caused by manual device movement and achieving sub-pixel alignment precision.
Solution Approach 2:
The patent creates multiple copies of the document view from different video frames, each containing feature points and descriptors. These copies are then processed through homography calculations to generate transformed versions that are precisely aligned and combined. The feature point descriptors serve as copies of local image characteristics that enable accurate matching and alignment across frames.
3Measurement precision
If feature point descriptors are computed from pixels surrounding feature points, then alignment accuracy improves, but computational complexity increases
Solution Approach 1:
The patent applies local quality by computing feature point descriptors only from pixels surrounding detected feature points, rather than processing the entire image. This localized approach concentrates computational resources on critical regions, achieving high matching precision for alignment while reducing overall computational complexity. The descriptor computation is performed selectively on small neighborhoods around key feature points.
Solution Approach 2:
The patent uses partial action by computing descriptors for only the most significant feature points detected in each video frame, rather than processing all pixels or all detected features. This selective computation provides sufficient alignment precision for document mosaicing while avoiding the excessive computational burden of processing every image region, achieving an optimal balance between accuracy and complexity.
4Ease of manufacture
If video frames are processed sequentially to assemble document image, then processing is simpler, but rendering speed decreases
Solution Approach 1:
The patent performs preliminary computations during video capture and initial processing phases, including feature point detection, descriptor computation, and homography matrix calculation for each frame. These pre-computed values are stored and then rapidly applied during the final assembly phase, enabling fast rendering of the complete document image without sacrificing processing simplicity. The complex operations are performed once in advance rather than repeatedly during assembly.
Data Source
AI summary
Aspects of the present disclosure propose techniques for reconstructing a document mosaic using video streams of the document. The streams provide information identifying a layout that relates sequential frames of a video to each other. Once the streams are captured using a mobile device, it is then possible to reconstruct a virtual view of the entire document as though it were taken with a single camera shot. The reconstructed virtual view of the document will be suitable as input to an optical character recognition engine, which can be used for translating the document.


