Document Tagging via Structural Block Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulty in quickly finding specific content within large electronic documents due to the lack of efficient methods for summarizing and structuring document information, leading to a need for automated tagging solutions that can determine suitable annotation positions and generate relevant tags.

Innovation Solution

A method and device for automatically generating structural tags in documents by acquiring and comparing structural information, retrieving content blocks, and annotating similar blocks with tags, allowing for precise tagging and annotation of document positions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging is performed by users reading and generalizing document content, then tag accuracy is improved, but time consumption and labor intensity increase significantly

Engineering Contradiction:
Improvetag accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables documents to tag themselves automatically by extracting structural information and comparing it with existing tagged documents. The document's own structure is used to identify similar blocks and apply appropriate tags without human intervention, making the tagging process self-service and eliminating manual labor while preserving accuracy through structural analysis

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system copies tag annotations from existing documents to new documents based on structural similarity. By comparing structural information blocks between documents, the system identifies similar blocks and copies the tags that were successfully applied to corresponding blocks in existing documents, thereby achieving accurate tagging without manual effort

Inventive Principle:
Principle #26Copying

2Loss of information

If manual tagging is performed for each document, then tag relevance is improved, but productivity decreases due to repetitive manual actions

Engineering Contradiction:
Improvetag relevanceVSAvoidtagging speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system retrieves tags from existing tagged documents and copies them to new documents based on structural similarity. This allows rapid tagging of multiple documents by reusing proven tags from existing documents, significantly increasing productivity while maintaining relevance through structure-based matching

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary tagging by automatically comparing structural information with existing documents and applying tags before any manual review is needed. This preliminary action handles the bulk of tagging work automatically, increasing productivity while allowing manual verification only when necessary to maintain tag relevance

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated tagging is implemented without structural analysis, then productivity increases, but tag accuracy and positioning precision decrease

Engineering Contradiction:
Improveautomation levelVSAvoidtag positioning accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system segments documents into structural information blocks (headers, paragraphs, lists, etc.) and processes each block independently for tagging. This segmentation allows automated processing of entire documents while maintaining precise positioning accuracy by analyzing and tagging each structural block based on its specific characteristics and similarity to blocks in existing documents

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8868556B2Method and device for tagging a document
Publication Date: 2014.10.21 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8868556B2 patent drawing
  • US8868556B2 patent drawing
  • US8868556B2 patent drawing

AI summary

There are disclosed a method and a device for tagging a document. The method includes the steps of acquiring structural information of the document, retrieving a content block list corresponding to a user-input tag, comparing blocks in the structural information with blocks in the content block list, to obtain similar blocks, and annotating the user-input tag at positions, which correspond to the similar blocks, in the document.