Automated Heading Classification for Digital Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document analytics systems are unable to accurately recognize and categorize higher-order structural features in digital documents, such as headings, due to limitations in document formats and the absence of explicit labeling, leading to inefficient and error-prone manual user intervention.

Innovation Solution

A document analysis system is developed to extract structural features from digital documents, identify headings, classify them into types, and generate a sectioned version of the document with a document directory for navigation, using feature extraction and machine learning models to determine heading likelihood and type.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual categorization of document features is implemented, then document structure recognition can be achieved, but user time consumption and error rates increase significantly

Engineering Contradiction:
Improvedocument structure recognition accuracyVSAvoiduser time consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The document analytics system automatically performs feature extraction, heading identification, and classification without requiring manual user intervention. The system extracts structural features, identifies headings based on visual attributes and position information, classifies them into types, and generates sectioned documents and directories autonomously, eliminating the need for users to manually categorize document features.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical categorization with an automated computational system using machine learning models and image processing techniques. The system uses a trained model to analyze visual attributes of structural features, determine heading likelihood, and automatically classify headings, substituting human manual work with automated intelligent processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated heading identification is implemented, then productivity increases, but measurement precision of heading classification may deteriorate

Engineering Contradiction:
Improvedocument processing speedVSAvoidheading classification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system employs a feedback mechanism where the machine learning model processes visual attributes and position information of structural features, generates heading likelihood scores, and uses this feedback to automatically classify headings. The model continuously refines its classification based on the analyzed features, enabling accurate automated heading identification while maintaining high processing speed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent utilizes parameter changes by analyzing multiple visual attributes (font size, bolding, italics, underline, position, spacing) and transforming them into a classification decision. The system changes the state of structural features from unclassified to classified by evaluating these parameters against trained model criteria, achieving both speed and accuracy in heading identification.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If detailed structural feature extraction is performed, then document analysis accuracy improves, but system complexity increases

Engineering Contradiction:
Improvedocument analysis accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The document analytics system segments the complex task of document analysis into distinct modules: feature extraction module that identifies structural features, heading identification module that determines headings from features, classification module that categorizes headings into types, and document generation module that creates sectioned documents and directories. This segmentation manages system complexity by dividing functions into specialized components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses an intermediary machine learning model that acts as a mediator between raw visual attributes and heading classification decisions. The model transforms complex visual feature data into simplified heading likelihood scores and classifications, reducing the complexity of direct analysis while maintaining high accuracy through the intermediary processing layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20210110153A1Heading Identification and Classification for a Digital Document
Publication Date: 2021.04.15 ADOBE INC
  • US20210110153A1 patent drawing
  • US20210110153A1 patent drawing
  • US20210110153A1 patent drawing

AI summary

Techniques described herein implement heading identification and classification for a digital document in a digital medium environment. A document analysis system is leveraged to extract structural features from a digital document, identify heading candidates from among the structural features, validate the headings candidates, and classify validated headings into different headings types. The classified headings are then utilized to generate a sectioned version of the digital document (“sectioned document”) that is divided into different sections based on the headings. Further, a document directory is generated that includes the headings and that enables navigation to different sections of the sectioned document.