Document Processing Device for Adaptive Character String Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques for extracting character strings from documents as bookmarks are inefficient in allowing users to correct conditions for extraction, leading to unintended character strings being extracted, and fail to adapt to different documentary forms effectively.
Innovation Solution
A document processing device and method that acquires document data, extracts character strings, derives features, displays them in a list format, and allows users to correct the format, with the system re-extracting strings based on the corrected format to ensure user-intended results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional automatic extraction techniques are used, then character strings can be extracted without user intervention, but the extracted strings may not satisfy user intentions and cannot adapt to different documentary forms
Solution Approach 1:
The system displays extracted character strings to users and receives feedback on whether each string should be extracted or not. This feedback is used to update the extraction condition, enabling the system to learn from user corrections and improve future extractions while maintaining automation.
Solution Approach 2:
The extraction condition is dynamically adjusted based on user feedback and document features. The system changes parameters such as extraction thresholds and criteria adaptively for different documentary forms, allowing it to maintain high automation while improving accuracy for various document types.
2Measurement precision
If extraction conditions are corrected manually for each document, then accuracy improves, but the operation becomes complex and time-consuming
Solution Approach 1:
The system performs self-learning by automatically updating extraction conditions based on user feedback. Instead of requiring manual correction of extraction parameters, the system adjusts itself, making the correction process simple and intuitive for users while significantly improving accuracy.
Solution Approach 2:
The system pre-adjusts extraction conditions based on the recognized documentary form before performing extraction. By preparing appropriate extraction parameters in advance for different document types, the system reduces the need for manual correction during the extraction process.
3Productivity
If extraction conditions are defined in advance, then processing speed is maintained, but the system cannot adapt to different documentary forms and user preferences
Solution Approach 1:
The extraction condition transitions from a static predefined state to a dynamic adaptive state. The system maintains fast processing by using predefined conditions as a baseline but dynamically adjusts these conditions based on document type recognition and user feedback, enabling both speed and adaptability.
Solution Approach 2:
The system creates a universal extraction framework that can handle multiple documentary forms. By recognizing document types and selecting or adjusting appropriate extraction conditions for each type, the system achieves versatility across different formats while maintaining efficient processing through automated selection.
Data Source
AI summary
A document processing device comprises: a document data acquiring part for acquiring document data; a character string extracting part for extracting character strings satisfying a predetermined condition for character string extraction from the document data acquired by the document data acquiring part; a format creating part for deriving the respective features of the character strings extracted by the character string extracting part, and for creating a format containing the derived features in the form of data; a display part on which the character strings extracted by the character string extracting part are displayed in a list form, and on which the format created by the format creating part is displayed; and a format correcting part for correcting the format displayed on the display part. The character string extracting part extracts character strings again to conform to the format corrected by the format correcting part.


