Topic Discovery HyperEngine Parallel Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for automatic topic discovery in large datasets, such as social media posts, are too slow for real-time applications, requiring hours, days, or even weeks to process, due to the high dimensionality of the computational problem involved.

Innovation Solution

A method utilizing a massively-parallel system architecture that includes a Harvester to collect and normalize data, a Publisher Discovery HyperEngine to refine publisher profiles, an Author Discovery HyperEngine to refine author profiles, and a Topic Discovery HyperEngine to generate statistical topic models using a trimmed lexicon, enabling near real-time topic discovery by clustering and filtering data streams with high-value information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional automatic topic discovery methods are used on large datasets, then comprehensive topic analysis is achieved, but processing time becomes excessively long (hours, days, or weeks)

Engineering Contradiction:
Improvetopic discovery accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the large corpus of electronic posts into smaller manageable chunks or batches that can be processed in parallel. By dividing the overall topic discovery task into multiple independent sub-tasks that can be executed simultaneously on different processors, the system maintains comprehensive analysis coverage while dramatically reducing total processing time from weeks to near real-time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of parallel processing by utilizing multiple processors to execute topic discovery algorithms simultaneously on different portions of the data. This dimensional expansion from sequential single-processor operation to parallel multi-processor operation enables the system to maintain analytical depth while achieving real-time performance.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If the lexicon size is increased to improve topic discovery quality, then better topic identification is achieved, but computational complexity and processing time increase

Engineering Contradiction:
Improvetopic identification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by using different lexicon sizes for different processing stages or different data segments. Rather than uniformly applying a large comprehensive lexicon to all data, the system can use smaller lexicons for initial filtering and larger lexicons for detailed analysis of specific topics, thereby reducing overall computational complexity while maintaining topic identification accuracy where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts the lexicon size parameter based on processing needs, data characteristics, and performance requirements. By changing the lexicon parameter adaptively rather than using a fixed large lexicon, the system reduces computational complexity while preserving the ability to achieve high topic identification accuracy when necessary.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If real-time processing is implemented to reduce processing time, then near real-time topic discovery is achieved, but processing depth and accuracy may be compromised

Engineering Contradiction:
Improveprocessing speedVSAvoidtopic analysis depth
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by pre-processing the data stream to identify and extract key features, terms, and patterns before the main topic discovery analysis. This preliminary filtering and feature extraction enables the subsequent real-time topic discovery to work with pre-processed, high-value information, maintaining analytical depth while achieving real-time processing speeds.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous topic discovery processing that operates uninterrupted on incoming data streams. Rather than batch processing that periodically suspends analysis, the system maintains continuous analytical action on the data flow, ensuring that topic discovery depth is preserved while achieving real-time responsiveness through ongoing parallel processing.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10599697B2Automatic topic discovery in streams of unstructured data
Publication Date: 2020.03.24 TARGET BRANDS INC
  • US10599697B2 patent drawing
  • US10599697B2 patent drawing
  • US10599697B2 patent drawing

AI summary

A method is provided for automatically discovering topics in electronic posts, such as social media posts. The method includes receiving a corpus that includes a plurality of electronic posts. The method further includes identifying a plurality of candidate terms within the corpus and selecting, as a trimmed lexicon, a subset of the plurality of candidate terms using predefined criteria. The method further includes clustering at least a subset of the plurality of electronic posts according to a plurality of clusters using the lexicon to produce a plurality of statistical topic models. The method further includes storing information corresponding to the statistical topic models.