Rough Set Morpheme Tagging Error Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for natural language processing fail to effectively detect and correct errors in corpora used for learning data, leading to increased time and costs due to manual production and correction of inconsistent mass corpora.

Innovation Solution

A device and method using rough sets to automatically detect and correct morpheme part-of-speech tagging corpus errors by generating attributes and calculating frequency counts for word phrases, applying a kernel to transform and analyze the corpus, and correcting errors based on statistical data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual production and correction of corpora is performed, then error detection and correction can be conducted, but time and costs increase significantly

Engineering Contradiction:
Improvecorpus error detection accuracyVSAvoidtime for manual correction
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-correction by automatically detecting corpus errors through rough set theory and correcting them without human intervention. The error detection apparatus analyzes the corpus, identifies inconsistencies, and corrects errors autonomously, eliminating the need for manual correction while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual correction process with an automated computational system. Rough set theory and statistical analysis algorithms substitute human operators, enabling automatic error detection and correction based on frequency counts and attribute analysis of word phrases.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual production and correction of corpora is performed, then error detection can be achieved, but costs increase

Engineering Contradiction:
Improvecorpus error detection accuracyVSAvoidcost of corpus correction
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system performs self-correction by automatically detecting corpus errors through rough set theory and correcting them without human intervention. The error detection apparatus analyzes the corpus, identifies inconsistencies, and corrects errors autonomously, eliminating the need for manual correction while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual correction process with an automated computational system. Rough set theory and statistical analysis algorithms substitute human operators, enabling automatic error detection and correction based on frequency counts and attribute analysis of word phrases.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If conventional error detection methods are used, then contextual or syntactic errors can be corrected, but corpus errors as learning data cannot be detected

Engineering Contradiction:
Improveerror correction capabilityVSAvoidapplicability to corpus learning data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent changes the detection parameters from contextual/syntactic error patterns to corpus-level attribute frequency analysis. By applying rough set theory and statistical methods to analyze word phrase attributes and their frequency counts across the corpus, the system detects errors specific to learning data that conventional methods cannot identify.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical manual correction process with an automated computational system. Rough set theory and statistical analysis algorithms substitute human operators, enabling automatic error detection and correction based on frequency counts and attribute analysis of word phrases.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11074406B2Device for automatically detecting morpheme part of speech tagging corpus error by using rough sets, and method therefor
Publication Date: 2021.07.27 CHANGWON NATIONAL UNIVERSITY INDUSTRY ACADEMY COOPERATION CORPS
  • US11074406B2 patent drawing
  • US11074406B2 patent drawing
  • US11074406B2 patent drawing

AI summary

A device for detecting a morpheme tagging corpus error, of the present invention, includes: an attribute generating unit for generating attributes for word phrases included in an input corpus, by using a kernel to which a rough set theory is applied; and an attribute statistics processing unit for generating part-of-speech tagging corpus error data through the calculation of attributes and frequency count for the same word phrases by counting attributes for the same word phrase among the word phrases, and thus the present invention can detect, quantify, and modify errors included in a corpus (learning data) required in learning for classifier generation and recognition for natural language processing.