Text Mining System for Reducing Analysis Cost
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text mining systems face significant time and labor costs when analyzing multiple pieces of data to identify differences, as they require exhaustive comparison of all data pairs and viewpoints, leading to increased analysis costs.
Innovation Solution
A text mining system that includes an analysis target data pair search unit, analysis viewpoint generation unit, positive example set identification unit, characteristic quantity calculation unit, and characteristic expression ranking unit, which identifies commonalities and differences in text data to prioritize analysis based on characteristic expression rank variations, reducing the need for exhaustive comparisons.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If exhaustive comparison of all analysis target data pairs is performed to identify differences, then analysis completeness is improved, but analysis cost increases remarkably
Solution Approach 1:
The patent segments the exhaustive comparison task into two stages: first, automated extraction of characteristic expressions and ranking from multiple text data sets; second, comparison of these extracted features rather than full-text comparison. This segmentation reduces the comparison scope while maintaining analysis completeness by focusing on discriminative features.
Solution Approach 2:
The system performs preliminary extraction and ranking of characteristic expressions from all text data sets before the comparison phase. By pre-processing and organizing the data into ranked characteristic expressions, the system prepares the information in advance, enabling efficient comparison without re-processing during the analysis phase.
2Measurement precision
If all analysis target data pairs and viewpoints are compared exhaustively, then difference detection accuracy is improved, but productivity deteriorates
Solution Approach 1:
The patent applies local quality by comparing only the characteristic expressions and their ranks rather than entire text data sets. This localized comparison focuses computational resources on the discriminative features that actually differ between data sets, maintaining detection accuracy while improving efficiency.
Solution Approach 2:
The system transforms the comparison task from comparing raw text data to comparing extracted parameters (characteristic expressions and their ranks). This parameter transformation simplifies the comparison operation and enables efficient identification of differences through rank variations.
3Measurement precision
If characteristic expression extraction is performed on all text data sets, then analysis thoroughness is improved, but device complexity increases
Solution Approach 1:
The characteristic expression extraction means is designed as a universal module that processes multiple text data sets using the same extraction and ranking logic. This multi-functional approach maintains analysis thoroughness while avoiding the need for separate processing systems for each data set, thereby controlling system complexity.
Data Source
AI summary
A text mining system including an analysis target search unit which judges whether a commonality in expressions among text data exists, an analysis viewpoint generation unit which generates an analysis viewpoint to extract an expression from the target data, a positive example set identification unit which identifies a positive example set including an expression matching the generated analysis viewpoint in the target data, a characteristic quantity calculation unit which calculates a characteristic quantity showing a degree of characterizing the positive example set of expressions in the target data, and a characteristic expression ranking unit which extracts expressions having the calculated characteristic quantity equal to or greater than a predetermined threshold as characteristic expressions and ranks the extracted characteristic expressions, and the target search unit extracts the analysis viewpoint among which a difference in ranks provided for the characteristic expressions is equal to or greater than a predetermined threshold.


