Text Mining Apparatus Using Confidence-Based Inherent Portion Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text data generated by computer processing, such as speech recognition, often contains errors, making it difficult to accurately discriminate and divide inherent and common portions from other text data, which hinders the practical implementation of cross-channel text mining.

Innovation Solution

A text mining apparatus and method that sets confidence for each text data piece and uses this confidence to extract inherent portions relative to other text data pieces, employing frequency calculations and mutual information to improve discrimination accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition is used to obtain text data, then the coverage of consumer requests is improved, but recognition errors cause imprecise feature word extraction

Engineering Contradiction:
Improvecoverage of consumer requestsVSAvoidfeature word extraction precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces call memo text data as an intermediary to verify and correct speech-recognized text data. The operator-created memos serve as a reference standard to identify and eliminate recognition errors, thereby maintaining high coverage while improving extraction precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses call memo text data as feedback to evaluate and correct the speech-recognized text data. By comparing the two data sources and using the memo data as a reference, the system identifies recognition errors and adjusts the extracted feature words accordingly.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If text mining is performed on both speech-recognized text data and call memo text data, then analysis comprehensiveness is improved, but difficulty in discriminating inherent portions increases

Engineering Contradiction:
Improveanalysis comprehensivenessVSAvoidinherent portion discrimination difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies local quality by treating the two text data sources differently based on their characteristics. Call memo text data is treated as the reference standard with higher reliability, while speech-recognized text data is treated as the data to be verified. This differential treatment simplifies the discrimination process by establishing a clear hierarchy between the two data sources.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If confidence is set for each text data piece, then discrimination accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvediscrimination accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameter of text data by associating confidence levels with each data piece. This parameter change enables the system to differentiate between high-confidence call memo data and lower-confidence speech-recognized data, improving discrimination accuracy without requiring complex structural modifications.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8751531B2Text mining apparatus, text mining method, and computer-readable recording medium
Publication Date: 2014.06.10 NEC CORP
  • US8751531B2 patent drawing
  • US8751531B2 patent drawing
  • US8751531B2 patent drawing

AI summary

A text mining apparatus, a text mining method, and a program are provided that accurately discriminate inherent portions of each of a plurality of text data pieces including a text data piece generated by computer processing.A text mining apparatus 1 to be used performs text mining using, as targets, a plurality of text data pieces including a text data piece generated by computer processing. Confidence is set for each of the text data pieces. The text mining apparatus 1 includes an inherent portion extraction unit 6 that extracts an inherent portion of each text data piece relative to another of the text data pieces, using the confidence set for each of the text data pieces.