Document Image Semantic Segmentation for Mobile Readability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional OCR techniques fail to accurately convert old or damaged documents into editable text, leading to suboptimal viewability and readability, especially on mobile devices with limited resources, and require cumbersome manual processes for customization.

Innovation Solution

A method that processes document images to identify semantically meaningful segments like tables of contents and page numbers, creating a digital representation that can be rendered differently based on client device parameters, using an optical character recognition algorithm and semantic component identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional OCR techniques are used to convert documents to text, then the process is automated, but the accuracy and readability of old or damaged documents deteriorates

Engineering Contradiction:
Improveautomation of document conversionVSAvoidaccuracy of text conversion
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent segments the document into distinct semantic components (text regions, tables, figures, headers, footers) and processes each segment separately using specialized OCR and recognition techniques. This allows conventional OCR to handle text segments accurately while alternative methods handle damaged or complex segments, thereby maintaining high automation while improving overall accuracy for old or damaged documents.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing quality levels to different parts of the document based on their semantic importance and difficulty. Critical text regions receive high-precision processing while less critical elements like headers or footers receive standardized processing. This local quality approach maintains automation while improving accuracy for difficult-to-process segments.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If conventional OCR techniques are used to replicate document appearance, then the output matches the original, but the readability and viewability on mobile devices deteriorates

Engineering Contradiction:
Improvereplication of document appearanceVSAvoidreadability on mobile devices
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent creates a dynamic document representation that can adapt its layout and formatting based on the target display device. The same semantic document structure is rendered differently for mobile devices (optimized for small screens and low network speeds) versus desktop devices. This dynamic adaptation maintains manufacturing precision for the source document while improving ease of operation for the target display environment.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes display parameters such as font size, line spacing, and layout configuration based on the detected device characteristics. For mobile devices, the system adjusts parameters to optimize readability on small screens and compensate for lower network transfer speeds, thereby improving ease of operation while preserving the original document's semantic integrity.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If manual techniques are used to create and customize document displays, then the document can be optimized for specific devices, but the time and effort required increases

Engineering Contradiction:
Improvecustomization of document displayVSAvoidtime for manual processing
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent implements self-service automation where the system automatically detects device characteristics, identifies semantic components, and generates optimized document representations without requiring manual intervention. The system serves itself by automatically adapting the document format based on detected device parameters, thereby eliminating the time-consuming manual customization process while maintaining optimal display quality.

Inventive Principle:
Principle #25Self-service

4Loss of information

If digital images of documents are transmitted over networks, then the complete visual information is preserved, but the transmission speed and memory requirements worsen

Engineering Contradiction:
Improvepreservation of visual informationVSAvoidnetwork transmission speed
Core Design Contradiction:
Loss of informationVSSpeed

Solution Approach 1:

The patent extracts only the essential semantic information from the document image (text content, structural relationships, semantic component identifiers) and transmits this condensed representation rather than the complete image data. This extraction approach preserves the critical information needed for display while dramatically reducing the data size and improving network transmission speed.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a compressed semantic copy of the document that captures the essential information in a format optimized for network transmission. This copy uses efficient data structures to represent semantic components, allowing quick transmission and reconstruction on the target device without requiring transfer of the original high-resolution image data.

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient and customizable display of documents on various devices by accurately capturing document semantics and formatting text and images for optimal readability, even on devices with limited capabilities.

Implementation Method 1

applies an optical character recognition algorithm to an image of a document to obtain a plurality of document segments, each document segment corresponding to a region of the image of the document and having associated recognized text

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS8254681B1Display of document image optimized for reading
Publication Date: 2012.08.28 GOOGLE LLC
  • US8254681B1 patent drawing
  • US8254681B1 patent drawing
  • US8254681B1 patent drawing

AI summary

Semantically meaningful segments of an image of a document, such as tables of contents, page numbers, footnotes, and the like, are identified. These segments form a model of the document image, which may then be rendered differently for different client devices. The rendering may be based on a display parameter provided by a client device, such as a display resolution of the client device, or a requested display format.