Document Tagging via Structural Block Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulty in quickly finding specific content within large electronic documents due to the lack of efficient methods for summarizing and structuring document information, leading to a need for automated tagging solutions that can determine suitable annotation positions and generate relevant tags.
Innovation Solution
A method and device for automatically generating structural tags in documents by acquiring and comparing structural information, retrieving content blocks, and annotating similar blocks with tags, allowing for precise tagging and annotation of document positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging is performed by users reading and generalizing document content, then tag accuracy is improved, but time consumption and labor intensity increase significantly
Solution Approach 1:
The system enables documents to tag themselves automatically by extracting structural information and comparing it with existing tagged documents. The document's own structure is used to identify similar blocks and apply appropriate tags without human intervention, making the tagging process self-service and eliminating manual labor while preserving accuracy through structural analysis
Solution Approach 2:
The system copies tag annotations from existing documents to new documents based on structural similarity. By comparing structural information blocks between documents, the system identifies similar blocks and copies the tags that were successfully applied to corresponding blocks in existing documents, thereby achieving accurate tagging without manual effort
2Loss of information
If manual tagging is performed for each document, then tag relevance is improved, but productivity decreases due to repetitive manual actions
Solution Approach 1:
The system retrieves tags from existing tagged documents and copies them to new documents based on structural similarity. This allows rapid tagging of multiple documents by reusing proven tags from existing documents, significantly increasing productivity while maintaining relevance through structure-based matching
Solution Approach 2:
The system performs preliminary tagging by automatically comparing structural information with existing documents and applying tags before any manual review is needed. This preliminary action handles the bulk of tagging work automatically, increasing productivity while allowing manual verification only when necessary to maintain tag relevance
3Productivity
If automated tagging is implemented without structural analysis, then productivity increases, but tag accuracy and positioning precision decrease
Solution Approach 1:
The system segments documents into structural information blocks (headers, paragraphs, lists, etc.) and processes each block independently for tagging. This segmentation allows automated processing of entire documents while maintaining precise positioning accuracy by analyzing and tagging each structural block based on its specific characteristics and similarity to blocks in existing documents
Data Source
AI summary
There are disclosed a method and a device for tagging a document. The method includes the steps of acquiring structural information of the document, retrieving a content block list corresponding to a user-input tag, comparing blocks in the structural information with blocks in the content block list, to obtain similar blocks, and annotating the user-input tag at positions, which correspond to the similar blocks, in the document.


