Hierarchical Word Topic Estimation with Parent Constraints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Latent topic estimation methods that handle mixtures of topics require significant processing time proportional to the number of topics, making them inefficient when the number of topics is large.

Innovation Solution

A word latent topic estimation device and method that utilize a hierarchical structure, where a higher-level constraint is created based on topic estimation results, allowing the estimation of latent topics to be performed efficiently by referencing the probability of assignment to parent topics as weights for lower-level topic estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If latent topic estimation methods handle mixtures of topics, then the accuracy of word topic probability estimation is improved, but the processing time increases proportionally to the number of topics

Engineering Contradiction:
Improveword topic probability estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the topic estimation process by introducing a higher-level constraint layer that divides the estimation into coarse-grained topic selection and fine-grained probability refinement. The higher-level constraint creation unit generates constraints that partition the search space, allowing the system to handle mixture topics accurately without proportionally increasing processing time for all topics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by creating higher-level constraints before performing detailed topic estimation. These constraints are generated in advance based on document characteristics and topic hierarchies, pre-filtering the topic space to reduce the computational burden during actual estimation while maintaining accurate mixture topic handling.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the number of topics is increased, then the ability to represent multiple topics in documents is improved, but the processing time becomes excessive

Engineering Contradiction:
Improvemulti-topic representation capabilityVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the large topic space into hierarchical levels with higher-level constraints that group related topics. This segmentation allows the system to maintain high adaptability for representing multiple diverse topics while improving productivity by processing topics in organized groups rather than as a monolithic set.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to the topic structure, organizing topics into multiple levels with parent-child relationships. This dimensional change allows the system to handle a large number of topics efficiently by navigating the hierarchical structure rather than processing all topics at a single level, thus maintaining versatility while improving processing efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9519633B2Word latent topic estimation device and word latent topic estimation method
Publication Date: 2016.12.13 NEC CORP
  • US9519633B2 patent drawing
  • US9519633B2 patent drawing
  • US9519633B2 patent drawing

AI summary

Provided are a word latent topic estimation device and a word latent topic estimation method which are capable of hierarchically performing processing and which are capable of rapidly estimating latent topics of a word while taking into consideration a mixed state of topics. The present invention is provided with: a document data addition unit (11) which inputs a document which includes one or more words; a level setting unit (12) which sets a number of topics at each level in accordance with a hierarchical structure of topics for hierarchically estimating latent topics of a word; a higher-level constraint creation unit (15) which, on the basis of results of topic estimation at a given level with regard to a word within the document, creates a higher-level constraint indicating an identifier of a topic for which there is a possibility of being assigned to the word and a probability of being assigned to the topic; and a higher-level-constraint-attached topic estimation unit (13) which, when estimating the probability of each word being assigned to each topic, refers to the higher-order constraint, uses the probability of being assigned to a parent topic at the higher level as a weight, and performs estimation processing to a lower-level topic.