Image Document Component Extraction for Flow Document Creation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current OCR technologies prioritize visual fidelity over preserving the flow when converting image documents to flow documents, resulting in unrecognizable section properties and section breaks at the end of each page.

Innovation Solution

The detection and extraction of text, paths, and images from image documents using optical character recognition (OCR) followed by binarization, allowing for the creation of a flow document that adapts to various reading experiences and provides editable documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current OCR technology is used to convert image documents to flow documents, then visual fidelity is improved, but document flow preservation deteriorates

Engineering Contradiction:
Improvevisual fidelityVSAvoiddocument flow
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the image document into distinct components (text regions, image regions, table regions, section breaks) using layout analysis. This segmentation allows the system to process and preserve each element's structural information separately, maintaining document flow while achieving accurate visual fidelity through targeted OCR and component reconstruction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary layout analysis and component detection before conducting OCR. By identifying text regions, section breaks, and structural elements in advance, the system preserves document flow information that would otherwise be lost during conversion, while still achieving high visual fidelity in the final flow document.

Inventive Principle:
Principle #10Preliminary action

2Shape

If section breaks are placed at the end of each page, then page structure is maintained, but section property recognition deteriorates

Engineering Contradiction:
Improvepage structureVSAvoidsection properties
Core Design Contradiction:
ShapeVSLoss of information

Solution Approach 1:

The patent uses feedback mechanisms where the layout analysis results inform the placement of section breaks and the preservation of section properties. The system analyzes the document structure, identifies meaningful section boundaries, and uses this information to guide the conversion process, ensuring both page structure maintenance and section property recognition.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies different processing strategies to different regions of the document. Rather than uniformly placing section breaks at page ends, the system identifies specific locations where section breaks should be preserved based on local document characteristics, maintaining both page structure and section property information through region-specific handling.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If multiple images are converted to flow document, then image content is preserved, but section detection complexity increases

Engineering Contradiction:
Improveimage contentVSAvoidsection detection
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the document into distinct regions including image regions, text regions, and table regions. By separating image detection from text processing and using region-based analysis, the system efficiently handles multiple images without significantly increasing overall detection complexity, as each region type is processed with specialized algorithms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a universal layout analysis framework that handles multiple component types (text, images, tables, section breaks) through a single integrated process. This multi-functional approach detects and processes various elements simultaneously, avoiding the need for separate complex detection systems for each component type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9355313B2Detecting and extracting image document components to create flow document
Publication Date: 2016.05.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9355313B2 patent drawing
  • US9355313B2 patent drawing
  • US9355313B2 patent drawing

AI summary

One or more components of an image document may be detected and extracted in order to create a flow document from the image document. Components of an image document may include text, one or more paths, and one or more images. The text may be detected using optical character recognition (OCR) and the image document may be binarized. The detected text may be extracted from the binarized image document to enable detection of the paths, which may then be extracted from the binarized image document to enable detection of the images. In some examples, the images, similar to the text and paths, may be extracted from the binarized image document. The extracted text, paths, and/or images may be stored in a data store, and may be retrieved in order to create a flow document that may provide better adaption to a variety of reading experiences and provide editable documents.