Document Processing Device for Adaptive Character String Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for extracting character strings from documents as bookmarks are inefficient in allowing users to correct conditions for extraction, leading to unintended character strings being extracted, and fail to adapt to different documentary forms effectively.

Innovation Solution

A document processing device and method that acquires document data, extracts character strings, derives features, displays them in a list format, and allows users to correct the format, with the system re-extracting strings based on the corrected format to ensure user-intended results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional automatic extraction techniques are used, then character strings can be extracted without user intervention, but the extracted strings may not satisfy user intentions and cannot adapt to different documentary forms

Engineering Contradiction:
Improveautomatic extractionVSAvoidextraction accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system displays extracted character strings to users and receives feedback on whether each string should be extracted or not. This feedback is used to update the extraction condition, enabling the system to learn from user corrections and improve future extractions while maintaining automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The extraction condition is dynamically adjusted based on user feedback and document features. The system changes parameters such as extraction thresholds and criteria adaptively for different documentary forms, allowing it to maintain high automation while improving accuracy for various document types.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If extraction conditions are corrected manually for each document, then accuracy improves, but the operation becomes complex and time-consuming

Engineering Contradiction:
Improveextraction accuracyVSAvoidcorrection operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-learning by automatically updating extraction conditions based on user feedback. Instead of requiring manual correction of extraction parameters, the system adjusts itself, making the correction process simple and intuitive for users while significantly improving accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-adjusts extraction conditions based on the recognized documentary form before performing extraction. By preparing appropriate extraction parameters in advance for different document types, the system reduces the need for manual correction during the extraction process.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If extraction conditions are defined in advance, then processing speed is maintained, but the system cannot adapt to different documentary forms and user preferences

Engineering Contradiction:
Improveprocessing speedVSAvoidadaptability to documentary forms
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The extraction condition transitions from a static predefined state to a dynamic adaptive state. The system maintains fast processing by using predefined conditions as a baseline but dynamically adjusts these conditions based on document type recognition and user feedback, enabling both speed and adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system creates a universal extraction framework that can handle multiple documentary forms. By recognizing document types and selecting or adjusting appropriate extraction conditions for each type, the system achieves versatility across different formats while maintaining efficient processing through automated selection.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8854635B2Document processing device, method, and recording medium for creating and correcting formats for extracting characters strings
Publication Date: 2014.10.07 KONICA MINOLTA BUSINESS TECH INC
  • US8854635B2 patent drawing
  • US8854635B2 patent drawing
  • US8854635B2 patent drawing

AI summary

A document processing device comprises: a document data acquiring part for acquiring document data; a character string extracting part for extracting character strings satisfying a predetermined condition for character string extraction from the document data acquired by the document data acquiring part; a format creating part for deriving the respective features of the character strings extracted by the character string extracting part, and for creating a format containing the derived features in the form of data; a display part on which the character strings extracted by the character string extracting part are displayed in a list form, and on which the format created by the format creating part is displayed; and a format correcting part for correcting the format displayed on the display part. The character string extracting part extracts character strings again to conform to the format corrected by the format correcting part.