Text Mining Apparatus Using Confidence-Based Inherent Portion Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text data generated by computer processing, such as speech recognition, often contains errors, making it difficult to accurately discriminate and divide inherent and common portions from other text data, which hinders the practical implementation of cross-channel text mining.
Innovation Solution
A text mining apparatus and method that sets confidence for each text data piece and uses this confidence to extract inherent portions relative to other text data pieces, employing frequency calculations and mutual information to improve discrimination accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is used to obtain text data, then the coverage of consumer requests is improved, but recognition errors cause imprecise feature word extraction
Solution Approach 1:
The patent introduces call memo text data as an intermediary to verify and correct speech-recognized text data. The operator-created memos serve as a reference standard to identify and eliminate recognition errors, thereby maintaining high coverage while improving extraction precision.
Solution Approach 2:
The system uses call memo text data as feedback to evaluate and correct the speech-recognized text data. By comparing the two data sources and using the memo data as a reference, the system identifies recognition errors and adjusts the extracted feature words accordingly.
2Adaptability or versatility
If text mining is performed on both speech-recognized text data and call memo text data, then analysis comprehensiveness is improved, but difficulty in discriminating inherent portions increases
Solution Approach 1:
The patent applies local quality by treating the two text data sources differently based on their characteristics. Call memo text data is treated as the reference standard with higher reliability, while speech-recognized text data is treated as the data to be verified. This differential treatment simplifies the discrimination process by establishing a clear hierarchy between the two data sources.
3Measurement precision
If confidence is set for each text data piece, then discrimination accuracy is improved, but device complexity increases
Solution Approach 1:
The patent changes the parameter of text data by associating confidence levels with each data piece. This parameter change enables the system to differentiate between high-confidence call memo data and lower-confidence speech-recognized data, improving discrimination accuracy without requiring complex structural modifications.
Data Source
AI summary
A text mining apparatus, a text mining method, and a program are provided that accurately discriminate inherent portions of each of a plurality of text data pieces including a text data piece generated by computer processing.A text mining apparatus 1 to be used performs text mining using, as targets, a plurality of text data pieces including a text data piece generated by computer processing. Confidence is set for each of the text data pieces. The text mining apparatus 1 includes an inherent portion extraction unit 6 that extracts an inherent portion of each text data piece relative to another of the text data pieces, using the confidence set for each of the text data pieces.


