Text Mining Method Correcting Feature Degree by Topic Relatedness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text mining technologies struggle to accurately analyze specific topics in text sets with varying degrees of involvement, leading to inclusion of unimportant elements and reduced accuracy due to overlapping topics and differing levels of content depth.
Innovation Solution
A text mining method that calculates a feature degree for each element, correcting it based on a topic relatedness degree to accurately identify distinctive elements within the text set by dividing the text into predetermined units and adjusting the feature degree according to the involvement in the analysis target topic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If text mining is performed on the entire text set including multiple topics, then the analysis covers all topics, but the accuracy for specific topic analysis deteriorates due to inclusion of unimportant elements from other topics
Solution Approach 1:
The patent segments the text set into multiple topic-specific subsets using topic analysis. Each topic subset contains only texts relevant to a specific topic, allowing separate text mining for each topic. This segmentation resolves the contradiction by enabling both comprehensive topic coverage (analyzing multiple topics) and high specific topic accuracy (excluding unrelated elements)
Solution Approach 2:
The patent applies different analysis approaches to different parts of the text set based on their topic classification. Each topic subset receives tailored text mining analysis appropriate to its specific topic characteristics. This local quality approach allows the system to maintain high accuracy for each specific topic while covering multiple topics overall
2Measurement precision
If text is divided into parts and only parts corresponding to the analysis target topic are analyzed, then the specific topic analysis accuracy improves, but the device complexity increases due to additional topic analysis step
Solution Approach 1:
The patent merges topic analysis functionality with text mining functionality into an integrated system. The topic analysis unit and text mining unit work together as a unified process, where topic analysis automatically guides the text mining process. This merging reduces device complexity compared to having completely separate systems, while maintaining high specific topic analysis accuracy
3Ease of operation
If all elements in the text set are treated equally in text mining, then the processing is simple, but the accuracy deteriorates due to varying degrees of involvement in different topics
Solution Approach 1:
The patent applies local quality by treating different elements differently based on their topic relevance. Elements within topic-specific subsets are processed with appropriate weighting and focus according to their specific topic context. This resolves the contradiction by maintaining processing simplicity through automated topic-based grouping while improving accuracy through differentiated element treatment
Data Source
AI summary
Disclosed are a text mining method, device, and program capable of performing text mining with a specific topic as an object with high precision. An element identification unit calculates a feature degree, which is an index for indicating a degree that within a text set of interest, which is a set of text that is to be analyzed, an element of the text appears. An output unit identifies distinctive elements within the text set of interest on the basis of the calculated feature degree and outputs the identified elements. The element identification unit corrects the feature degree on the basis of a topic relatedness degree, which is a value indicating a degree related to a topic of analysis, which is a topic for which each text portion of the text being analyzed has been partitioned into predetermined units that are to be analyzed.


