Rough Set Morpheme Tagging Error Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for natural language processing fail to effectively detect and correct errors in corpora used for learning data, leading to increased time and costs due to manual production and correction of inconsistent mass corpora.
Innovation Solution
A device and method using rough sets to automatically detect and correct morpheme part-of-speech tagging corpus errors by generating attributes and calculating frequency counts for word phrases, applying a kernel to transform and analyze the corpus, and correcting errors based on statistical data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual production and correction of corpora is performed, then error detection and correction can be conducted, but time and costs increase significantly
Solution Approach 1:
The system performs self-correction by automatically detecting corpus errors through rough set theory and correcting them without human intervention. The error detection apparatus analyzes the corpus, identifies inconsistencies, and corrects errors autonomously, eliminating the need for manual correction while maintaining high reliability.
Solution Approach 2:
The patent replaces the mechanical manual correction process with an automated computational system. Rough set theory and statistical analysis algorithms substitute human operators, enabling automatic error detection and correction based on frequency counts and attribute analysis of word phrases.
2Reliability
If manual production and correction of corpora is performed, then error detection can be achieved, but costs increase
Solution Approach 1:
The system performs self-correction by automatically detecting corpus errors through rough set theory and correcting them without human intervention. The error detection apparatus analyzes the corpus, identifies inconsistencies, and corrects errors autonomously, eliminating the need for manual correction while maintaining high reliability.
Solution Approach 2:
The patent replaces the mechanical manual correction process with an automated computational system. Rough set theory and statistical analysis algorithms substitute human operators, enabling automatic error detection and correction based on frequency counts and attribute analysis of word phrases.
3Reliability
If conventional error detection methods are used, then contextual or syntactic errors can be corrected, but corpus errors as learning data cannot be detected
Solution Approach 1:
The patent changes the detection parameters from contextual/syntactic error patterns to corpus-level attribute frequency analysis. By applying rough set theory and statistical methods to analyze word phrase attributes and their frequency counts across the corpus, the system detects errors specific to learning data that conventional methods cannot identify.
Solution Approach 2:
The patent replaces the mechanical manual correction process with an automated computational system. Rough set theory and statistical analysis algorithms substitute human operators, enabling automatic error detection and correction based on frequency counts and attribute analysis of word phrases.
Data Source
AI summary
A device for detecting a morpheme tagging corpus error, of the present invention, includes: an attribute generating unit for generating attributes for word phrases included in an input corpus, by using a kernel to which a rough set theory is applied; and an attribute statistics processing unit for generating part-of-speech tagging corpus error data through the calculation of attributes and frequency count for the same word phrases by counting attributes for the same word phrase among the word phrases, and thus the present invention can detect, quantify, and modify errors included in a corpus (learning data) required in learning for classifier generation and recognition for natural language processing.


