Hierarchical Clustering for Textual Data Theme Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting sales information from textual data, such as transcribed calls and emails, are inefficient due to limited accuracy in keyword identification and lack of understanding of themes, leading to incomplete analysis and manual, time-consuming processes that fail to uncover customer and market interests effectively.
Innovation Solution
A method using trained clustering and naming models to discover and aggregate themes in textual data, applying hierarchical clustering to identify clusters based on meaning, generating names for these clusters, and analyzing distribution metrics to provide notifications on significant changes, thereby facilitating objective and efficient analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If keyword search based on predefined dictionary is used, then identification speed is improved, but identification accuracy deteriorates
Solution Approach 1:
The patent replaces the mechanical keyword-matching system with a machine learning-based semantic analysis system. Instead of using predefined dictionaries and simple string matching, the system employs trained models to understand the meaning and context of textual data, thereby maintaining high processing speed while significantly improving identification accuracy.
Solution Approach 2:
The patent changes the fundamental parameter of identification from exact keyword matching to semantic similarity measurement. By transforming the approach from discrete keyword comparison to continuous semantic space analysis, the system achieves both speed and accuracy through vector-based representations and similarity calculations.
2Loss of information
If manual analysis is used to uncover insights, then analysis depth is improved, but time consumption deteriorates
Solution Approach 1:
The patent enables the system to perform automated thematic analysis without requiring manual intervention. The trained clustering and naming models independently process large volumes of textual data, identify patterns, and generate insights automatically, eliminating the need for human analysts while maintaining comprehensive analysis depth.
Solution Approach 2:
The patent accelerates the analysis process by employing powerful machine learning models that can process and understand semantic relationships in textual data much faster than human analysts. The system's computational power acts as an accelerator, enabling deep analysis of vast datasets in minimal time.
3Reliability
If thematic analysis without user input is implemented, then objectivity is improved, but ability to discover unknown subjects deteriorates
Solution Approach 1:
The patent implements a dynamic system where the clustering models can adapt to different datasets and domains. The unsupervised learning approach allows the system to automatically adjust to new types of textual data and discover emerging themes without predefined constraints, maintaining both objectivity and adaptability.
Solution Approach 2:
The patent employs pre-trained models that have been trained on diverse datasets beforehand. These models possess pre-acquired knowledge of language patterns and semantic relationships, enabling them to objectively analyze new data while simultaneously discovering unknown subjects through their generalized understanding.
Data Source
AI summary
A system and method for discovering and aggregating themes. The method includes applying a trained clustering model to a plurality of textual data, wherein the trained clustering model determines at least one cluster of textual data based on a meaning of the textual data, wherein textual data of the at least one cluster is a portion of the plurality of textual data; generating a name, using a trained naming model, for each of the at least one cluster, wherein the generated name indicates a theme that represents the meaning of the textual data of the at least one cluster; analyzing the at least one cluster to determine a distribution metric of the at least one cluster; and generating a notification based on the determined distribution metric and the respective at least one cluster.


