Inferring Section Headings in Electronic Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic documents such as PDFs and OOXMLs often lack explicit identification of sections or section headings, making it difficult for users to search or retrieve information from specific sections, as the semantic information implied by the author is not specified using computer-recognizable information.

Innovation Solution

A method and system that infers section headings by generating a list of candidate headings based on statistical distributions of point sizes, iteratively identifying adjacent and child chain fragments, and embedding inferred section heading information into the document using OOXML tags or similar standards, allowing for computer-recognizable section identification and enhanced search functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If section headings are not explicitly identified in electronic documents, then the document format remains simple and compatible with various file standards, but the ability to search and retrieve information from specific sections is impaired

Engineering Contradiction:
Improvesearch and retrieve informationVSAvoidsemantic information identification
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent replaces manual or explicit semantic markup with automated computational analysis. A machine learning model analyzes document structure, text patterns, and formatting to automatically identify and tag section headings, converting an information loss problem into a computational pattern recognition task.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary processing layer between the raw document content and the search functionality. This intermediary system automatically generates section heading identifiers by analyzing document patterns, enabling search capabilities without requiring explicit markup in the original document.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If automated section heading identification is implemented, then search functionality is enhanced, but the complexity of document processing increases

Engineering Contradiction:
Improveinformation retrievalVSAvoiddocument processing system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent enables the document processing system to automatically identify and tag its own section headings without requiring external manual intervention. The machine learning model analyzes the document's inherent patterns and self-generates the necessary structural identifiers, making the system self-sufficient.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the document processing approach by changing parameters from explicit markup requirements to pattern-based automatic identification. The system uses statistical analysis of text patterns, formatting parameters, and structural features to infer section headings, converting a complex markup task into parameter-based pattern recognition.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If explicit section headings are added to electronic documents, then information retrieval is improved, but the document format compatibility and simplicity are reduced

Engineering Contradiction:
Improveinformation retrieval efficiencyVSAvoiddocument structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary automatic analysis of document structure before information retrieval operations. By pre-identifying and tagging section headings through pattern analysis, the system prepares the document structure in advance, enabling efficient search and retrieval without requiring manual preprocessing or complex formatting changes.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11494555B2Identifying section headings in a document
Publication Date: 2022.11.08 KONICA MINOLTA BUSINESS SOLUTIONS USA INC
  • US11494555B2 patent drawing
  • US11494555B2 patent drawing
  • US11494555B2 patent drawing

AI summary

A method, non-transitory computer readable medium, and system for inferring certain texts as stylized section headings in an electronic document (ED). Stylized section headings are section headings that have unique styling distinct from the body of text below each stylized heading. In particular, the stylized section headings are identified based on styling information in the ED. Identifying stylized section headings includes grouping candidate headings based on identification of dominant styling, locating high level fragments, and repeatedly locating nested fragments from within higher level fragments. The ED may or may not include explicitly identified headings in the document.