Document Analysis System Structural Feature Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document analytics systems are unable to recognize and categorize higher-order features in digital documents, such as bulleted lists, tables, and check boxes, which prevents them from generating accurately reformatted versions that preserve the semantic integrity of these features, especially when documents are adapted for display on different form factor devices.

Innovation Solution

A document analysis system employs machine learning techniques, using a character analysis model and a classification model to extract and classify structural features by feature type, generating modifiable versions of digital documents that preserve their semantic context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional OCR techniques are used to convert image-based text, then text recognition is achieved, but higher-order structural features cannot be recognized or categorized

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidstructural feature information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the document analysis process into multiple specialized components: OCR for text recognition, separate structural feature detection for identifying higher-order elements, and categorization modules for different feature types. This segmentation allows each component to specialize in its function while preserving both text accuracy and structural information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary data structures and processing layers between OCR output and final document representation. These intermediaries capture structural feature information that would otherwise be lost, acting as mediators that preserve both textual and structural characteristics of the document.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual approaches are used to reformat digital documents, then structural integrity can be preserved, but the process is infeasible for large volumes of documents

Engineering Contradiction:
Improvestructural integrity preservationVSAvoiddocument processing volume
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements automated document analysis and reformatting systems that perform structural feature recognition and document restructuring without human intervention. The system serves itself by using machine learning models to automatically identify, categorize, and preserve structural features while reformating documents for different devices, eliminating the need for manual processing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameters of document processing by transitioning from manual to automated methods, and from simple text extraction to comprehensive structural feature analysis. This allows the system to handle large volumes of documents while maintaining structural integrity through intelligent automated classification and reformatting.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If conventional document analytics systems are used, then digitization is achieved, but accurate reformatting that preserves semantic integrity cannot be performed

Engineering Contradiction:
Improvedigitization capabilityVSAvoidreformatting accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent performs preliminary analysis and categorization of structural features during the digitization process itself. By identifying and classifying higher-order features before reformatting operations, the system ensures that semantic integrity is preserved from the outset, rather than attempting to preserve it during subsequent reformatting steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent adds a new dimension to document analysis by incorporating structural feature recognition and categorization alongside traditional text extraction. This multi-dimensional approach captures both textual content and structural characteristics, enabling accurate reformatting that preserves semantic integrity across different device form factors.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11003862B2Classifying structural features of a digital document by feature type using machine learning
Publication Date: 2021.05.11 ADOBE INC
  • US11003862B2 patent drawing
  • US11003862B2 patent drawing
  • US11003862B2 patent drawing

AI summary

Classifying structural features of a digital document by feature type using machine learning is leveraged in a digital medium environment. A document analysis system is leveraged to extract structural features from digital documents, and to classifying the structural features by respective feature types. To do this, the document analysis system employs a character analysis model and a classification model. The character analysis model takes text content from a digital document and generates text vectors that represent the text content. A vector sequence is generated based on the text vectors and position information for structural features of the digital document, and the classification model processes the vector sequence to classify the structural features into different feature types. The document analysis system can generate a modifiable version of the digital document that enables its structural features to be modified based on their respective feature types.