Context-Sensitive Spelling Error Test Document Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional spelling error correction test techniques are inefficient and costly, relying on manually built test documents that do not comprehensively cover linguistic errors, limiting the reliability and scope of performance measurement for correction systems.

Innovation Solution

A system and method for generating context-sensitive test documents using statistical data from large-scale corpora, automatically creating error-intensive documents by extracting actual human errors, which are then used to measure the performance of spelling error correction systems, reducing costs and improving reliability through adaptive error handling based on surrounding context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If test documents are built manually by experts, then the quality and accuracy of test documents improve, but the cost and time consumption increase significantly

Engineering Contradiction:
Improvetest document qualityVSAvoidtest document preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates synthetic test documents by copying and transforming real error patterns from actual corpora. Instead of manually creating test cases, the system extracts error types from real documents and generates synthetic test documents that replicate these error patterns, achieving high test quality without manual effort

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating its own test documents using algorithms that analyze error patterns and create test cases without human intervention. The automated generation process eliminates the need for expert manual creation while maintaining test document quality through systematic error pattern replication

Inventive Principle:
Principle #25Self-service

2Reliability

If the number of test documents is increased to cover more linguistic errors, then the measurement reliability improves, but the cost and time for building test documents increase

Engineering Contradiction:
Improveerror coverage reliabilityVSAvoidtest document generation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent creates a universal test document generation system that can produce diverse test documents covering multiple error types from a single automated process. The system extracts general error patterns from corpora and applies them across numerous test documents, enabling comprehensive error coverage without proportional increases in manual work

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary action by pre-extracting and cataloging error patterns from large corpora before test document generation. This preliminary analysis creates a reusable error pattern library that accelerates subsequent test document creation, allowing rapid generation of numerous test documents with consistent quality

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If context-sensitive error correction is implemented, then the accuracy of error correction improves, but the complexity of the correction system increases

Engineering Contradiction:
Improvespelling error correction accuracyVSAvoidcorrection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by analyzing the specific context surrounding each error word and tailoring correction decisions to that local context. Instead of applying uniform correction rules, the system evaluates surrounding words, grammar, and semantics to determine the most appropriate correction for each specific error instance

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system introduces an intermediary context analysis layer between error detection and correction. This intermediary layer examines surrounding text, identifies error patterns, and mediates the correction process by selecting appropriate corrections based on contextual clues, achieving high accuracy without requiring overly complex direct correction mechanisms

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If actual human errors from large corpora are used for test generation, then the realism and effectiveness of tests improve, but the data processing complexity increases

Engineering Contradiction:
Improvetest realismVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies extraction by systematically extracting error patterns from large corpora and isolating them as reusable test components. The system identifies and extracts specific error types, error contexts, and correction patterns from raw corpus data, separating the essential error characteristics from the surrounding text to create concise test cases

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates simplified copies of actual error instances from corpora. Instead of processing entire raw documents, it extracts and copies only the essential error patterns and contextual information needed for testing, reducing data complexity while maintaining test realism through faithful replication of actual error characteristics

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11429785B2System and method for generating test document for context sensitive spelling error correction
Publication Date: 2022.08.30 PUSAN NAT UNIV IND UNIV COOPERATION FOUND
  • US11429785B2 patent drawing
  • US11429785B2 patent drawing

AI summary

Disclosed is a system for generating test documents for context-sensitive spelling error correction. The system includes: an input unit inputting an error-free document for generating an error document; an error target word segment test unit checking possibility of an error in a word segment by sequentially examining word segments of the entire sentences in the document input through the input unit and searching for a candidate word appearing at the corresponding position together with surrounding context; an error word candidate selection unit selecting error word candidates among candidate words found at the corresponding position by considering edit distances to a correct word and keyboard typographical errors; and an error word determination and presentation unit calculating probabilities of an error word candidate and its surrounding context and determining an error word of the highest priority as a final error word.