Title Inference System for Electronic Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic documents often lack explicit title identification, making it difficult for users to search for titles within the documents, even though title texts may be more memorable.

Innovation Solution

A method and system that process electronic documents to infer titles by generating a mark-up version with text-styling and text-layout attributes, calculating relative weight scores, and determining title confidence scores based on styling, layout, and content information to identify potential titles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If titles are not explicitly labeled or tagged in electronic documents, then document structure remains simple and processing is easier, but users cannot easily search for or identify titles within the document

Engineering Contradiction:
Improvetitle search capabilityVSAvoidtitle identification system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system automatically infers titles by analyzing text styling attributes (bold, italic, font size), layout attributes (position, spacing), and content characteristics without requiring manual labeling or tagging. The document structure itself provides the information needed for title identification through its inherent formatting and arrangement properties

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual title labeling mechanisms with an automated computational system that uses pattern recognition and statistical analysis to infer titles. The system processes text attributes and calculates confidence scores to automatically identify titles, eliminating the need for explicit metadata tags or manual annotations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If the system analyzes multiple text-styling and text-layout attributes to infer titles, then title identification accuracy improves, but processing time and computational complexity increase

Engineering Contradiction:
Improvetitle confidence score accuracyVSAvoiddocument processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system assigns different weights to different text-styling and text-layout attributes based on their relevance to title identification. By dynamically adjusting the importance of various attributes (such as giving higher weight to bold formatting and larger font sizes), the system optimizes the balance between processing comprehensive attributes and maintaining fast processing speeds

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system processes only the necessary text attributes required for reliable title inference rather than analyzing every possible document property. By selectively processing key styling and layout attributes that are most indicative of titles, the system achieves accurate title identification without the computational overhead of complete document analysis

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10572587B2Title inferencer
Publication Date: 2020.02.25 KONICA MINOLTA SYSTEMS LABORATORY INC
  • US10572587B2 patent drawing
  • US10572587B2 patent drawing
  • US10572587B2 patent drawing

AI summary

A method for processing an electronic document (ED) to infer titles in the ED is provided. The method includes: generating a mark-up version of the ED comprising text-styling attributes, text-layout attributes, and text content information of characters included in the ED; generating statistical information of the text-styling and text-layout attributes; calculating, for each text-styling and text-layout attribute, a relative weight score; calculating, for each paragraph in the ED: a styling criteria score and a layout criteria score based on the statistical information and the relative weight scores; a text content score based on the text content information; and a title confidence score based on the styling criteria score, the layout criteria score, and the text content score; and generating a metadata for the ED that includes the title confidence score for each paragraph for use in inferring the titles in the ED.