Text Mining Graph Node Merging for Semantic Concept Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text mining devices fail to properly extract characteristic structures when dealing with texts containing multiple words representing the same concept or semantically associated words, leading to incorrect identification and separation of concepts.

Innovation Solution

A data processing device with an association node extraction unit and an association node joint unit that transforms graphs by joining semantically associated nodes, allowing for the extraction of characteristic structures that represent identical or semantically associated concepts as a single entity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional text mining device processes texts with multiple words representing identical concepts, then it separates them into different characteristic structures, but this leads to incorrect identification and failure to extract proper characteristic structures

Engineering Contradiction:
Improvecharacteristic structure extraction accuracyVSAvoidconcept identification correctness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent merges semantically associated words (anaphoric pronouns, zero pronouns, and their antecedents) into unified nodes in the graph structure. This allows the text mining device to recognize that different word representations refer to the same concept, enabling correct extraction of characteristic structures that accurately represent the intended meaning rather than treating each word variant as a separate concept.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If the text mining device uses traditional parsing to create sentence structures, then it can analyze text structure, but it cannot identify cases where single words and multiple words describe the same concept

Engineering Contradiction:
Improvehandling of word variationsVSAvoidconcept equivalence detection
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces semantic association relationships as intermediary connections between nodes representing semantically associated words. This intermediary mechanism enables the system to recognize conceptual equivalence between single words and multiple words without requiring complex analysis, allowing the text mining device to adapt to different word representations while maintaining precise concept identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If the device extracts characteristic structures from parsed sentence structures, then it can identify frequent patterns, but it fails when texts use different word representations for the same concept

Engineering Contradiction:
Improvetext mining processing efficiencyVSAvoidcharacteristic structure extraction accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary processing by creating semantic association relationships between nodes before extracting characteristic structures. This preliminary action of establishing semantic connections ensures that when the extraction process runs, it can correctly identify and group semantically associated words, maintaining both processing efficiency and extraction accuracy even when texts use different word representations for the same concept.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8775158B2Data processing device, data processing method, and data processing program
Publication Date: 2014.07.08 NEC CORP
  • US8775158B2 patent drawing
  • US8775158B2 patent drawing
  • US8775158B2 patent drawing

AI summary

[PROBLEMS] To provide a data processing device such as a text mining device capable of extracting characteristic structures properly even in case a plurality of words indicating identical contents or a plurality of words semantically associated are contained in input data. [MEANS FOR SOLVING PROBLEMS] Association node extraction unit (22) of a text mining device (10) extracts association nodes containing semantically associated words from a graph obtained as a result of syntax analysis. Association node joint unit (23) transforms the graph by joint of a part of or a whole of the association nodes. Characteristic structure extraction unit (24) extracts a characteristic structure from the graph transformed by the association node joint unit.