Contract Document Normalization for Spelling-Variant Keywords
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional information processing devices struggle with handling keywords with spelling inconsistencies, making it difficult to manage relevant contract details uniformly.
Innovation Solution
A document processing method that extracts character strings with positional information, normalizes them, and displays the normalized information alongside the original document, enabling unified management of contract content despite spelling inconsistencies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If keyword detection is performed on contract text, then important parts can be recognized, but spelling inconsistencies cause inability to handle and manage relevant portions uniformly
Solution Approach 1:
The patent applies parameter changes by transforming the extracted character strings through normalization processes that standardize spelling variations. The system changes the textual parameters by applying correction rules, reference dictionary matching, and consistency transformations to convert inconsistent keyword spellings into unified forms, thereby resolving the contradiction between detection accuracy and adaptability to spelling variations
Solution Approach 2:
The patent introduces an intermediary normalization process between keyword detection and management. This intermediary step includes extraction means for obtaining character strings, normalization means for standardizing them, and management means for handling the normalized results. The intermediary processing layer mediates between raw extracted keywords with spelling inconsistencies and the final unified management system
2Loss of information
If extracted information is displayed without normalization, then original content is preserved, but spelling inconsistencies make unified management difficult
Solution Approach 1:
The patent applies parameter changes by transforming the extracted character strings through normalization processes that standardize spelling variations while preserving the ability to trace back to original positions. The system changes textual parameters through correction rules and reference matching, converting inconsistent spellings into unified forms for easier management
Solution Approach 2:
The patent segments the information processing into distinct functional modules: extraction means for obtaining character strings with positional information, normalization means for standardizing the extracted content, and management means for handling normalized information. This segmentation allows original content to be preserved in the extraction stage while enabling unified management after normalization
3Measurement precision
If keyword detection is implemented, then contract importance can be identified, but spelling variations prevent effective management of relevant portions
Solution Approach 1:
The patent applies parameter changes by transforming extracted keywords through normalization that standardizes spelling variations. The system changes textual parameters by applying correction rules, reference dictionary matching, and consistency transformations to convert inconsistent keyword spellings into unified forms, thereby improving reliable handling while maintaining detection precision
Solution Approach 2:
The patent introduces an intermediary normalization process between keyword detection and management. This intermediary includes extraction means for obtaining character strings, normalization means for standardizing them with positional tracking, and management means for handling normalized results. The intermediary processing layer ensures both precise detection and consistent management
Data Source
AI summary
A document processing method comprising: obtaining a character string indicating a content of a document extracting from document information; and obtaining a normalized extracted information by normalizing the character string information in the document information.


