Text Mining Method Correcting Feature Degree by Topic Relatedness

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text mining technologies struggle to accurately analyze specific topics in text sets with varying degrees of involvement, leading to inclusion of unimportant elements and reduced accuracy due to overlapping topics and differing levels of content depth.

Innovation Solution

A text mining method that calculates a feature degree for each element, correcting it based on a topic relatedness degree to accurately identify distinctive elements within the text set by dividing the text into predetermined units and adjusting the feature degree according to the involvement in the analysis target topic.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text mining is performed on the entire text set including multiple topics, then the analysis covers all topics, but the accuracy for specific topic analysis deteriorates due to inclusion of unimportant elements from other topics

Engineering Contradiction:
Improvetopic coverageVSAvoidspecific topic analysis accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the text set into multiple topic-specific subsets using topic analysis. Each topic subset contains only texts relevant to a specific topic, allowing separate text mining for each topic. This segmentation resolves the contradiction by enabling both comprehensive topic coverage (analyzing multiple topics) and high specific topic accuracy (excluding unrelated elements)

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different analysis approaches to different parts of the text set based on their topic classification. Each topic subset receives tailored text mining analysis appropriate to its specific topic characteristics. This local quality approach allows the system to maintain high accuracy for each specific topic while covering multiple topics overall

Inventive Principle:
Principle #3Local quality

2Measurement precision

If text is divided into parts and only parts corresponding to the analysis target topic are analyzed, then the specific topic analysis accuracy improves, but the device complexity increases due to additional topic analysis step

Engineering Contradiction:
Improvespecific topic analysis accuracyVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges topic analysis functionality with text mining functionality into an integrated system. The topic analysis unit and text mining unit work together as a unified process, where topic analysis automatically guides the text mining process. This merging reduces device complexity compared to having completely separate systems, while maintaining high specific topic analysis accuracy

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If all elements in the text set are treated equally in text mining, then the processing is simple, but the accuracy deteriorates due to varying degrees of involvement in different topics

Engineering Contradiction:
Improveprocessing simplicityVSAvoidelement identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent applies local quality by treating different elements differently based on their topic relevance. Elements within topic-specific subsets are processed with appropriate weighting and focus according to their specific topic context. This resolves the contradiction by maintaining processing simplicity through automated topic-based grouping while improving accuracy through differentiated element treatment

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9135326B2Text mining method, text mining device and text mining program
Publication Date: 2015.09.15 NEC CORP
  • US9135326B2 patent drawing
  • US9135326B2 patent drawing
  • US9135326B2 patent drawing

AI summary

Disclosed are a text mining method, device, and program capable of performing text mining with a specific topic as an object with high precision. An element identification unit calculates a feature degree, which is an index for indicating a degree that within a text set of interest, which is a set of text that is to be analyzed, an element of the text appears. An output unit identifies distinctive elements within the text set of interest on the basis of the calculated feature degree and outputs the identified elements. The element identification unit corrects the feature degree on the basis of a topic relatedness degree, which is a value indicating a degree related to a topic of analysis, which is a topic for which each text portion of the text being analyzed has been partitioned into predetermined units that are to be analyzed.