Pitman-Yor Topic Model Pre-Seeded Keyword Groups
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional admixture models for natural language processing are limited by their inability to utilize labeled training data or prior knowledge of themes, requiring large amounts of unlabelled data and extensive human intervention for interpretation, making them inefficient for classifying phone conversations.
Innovation Solution
A classification model based on a hierarchical Pitman-Yor process that incorporates pre-seeded topics and uses unlabeled data for training, allowing for the development of a predictive and interpretable model capable of leveraging prior knowledge, with a Bayesian Belief Network structure that enables efficient training and interpretation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional admixture models are used for topic modeling, then the model can operate without labeled training data, but the model requires large amounts of unlabeled data and extensive human intervention for interpretation
Solution Approach 1:
The patent applies preliminary action by pre-seeding the topic model with keyword groups before training. These keyword groups serve as initial topic representations that guide the model learning process, allowing the model to leverage prior knowledge and reduce the amount of unlabeled data needed for effective topic discovery and interpretation.
2Extent of automation
If conventional admixture models are used for topic modeling, then the model can operate without labeled training data, but extensive human intervention is required to interpret output and ascribe meaning to topics
Solution Approach 1:
The patent applies preliminary action by pre-seeding the topic model with keyword groups before training. These keyword groups serve as initial topic representations that guide the model learning process, allowing the model to leverage prior knowledge and reduce the amount of unlabeled data needed for effective topic discovery and interpretation.
Solution Approach 2:
The patent implements feedback through active learning loops where the model's topic assignments are evaluated and used to refine keyword groupings. This iterative feedback process automatically improves topic interpretation quality over time, reducing the need for manual human intervention while maintaining high interpretability.
3Reliability
If a highly predictive model is developed using conventional techniques, then the model can achieve good classification accuracy, but the model lacks interpretability and cannot be audited for fairness or correctness
Solution Approach 1:
The patent applies segmentation by separating the model into distinct interpretable components: topic models that identify thematic elements, keyword groups that represent meaningful concepts, and classification layers that combine these elements. This segmented architecture maintains predictive accuracy while enabling auditability of each component's contribution to final predictions.
Solution Approach 2:
The patent introduces topic representations as intermediary elements between raw input data and final classifications. These topics serve as interpretable mediators that bridge the gap between complex predictive processing and human-understandable reasoning, allowing auditors to trace decision pathways through meaningful thematic categories.
4Reliability
If labeled training data is used to improve model accuracy, then the model can achieve better classification performance, but the cost of curating labeled data becomes prohibitively expensive
Solution Approach 1:
The patent applies preliminary action by pre-seeding the topic model with keyword groups before training. These keyword groups serve as initial topic representations that guide the model learning process, allowing the model to leverage prior knowledge and reduce the amount of unlabeled data needed for effective topic discovery and interpretation.
Solution Approach 2:
The patent implements feedback through active learning loops where the model's topic assignments are evaluated and used to refine keyword groupings. This iterative feedback process automatically improves topic interpretation quality over time, reducing the need for manual human intervention while maintaining high interpretability.
Data Source
AI summary
In one embodiment, the disclosed technology involves: digitally generating and storing a machine learning statistical topic model in computer memory, the topic model being programmed to model call transcript data representing words spoken on a call as a function of one or more topics of a set of topics that includes pre-seeded topics and non-pre-seeded topics; programmatically pre-seeding the topic model with a set of keyword groups; programmatically training the topic model using unlabeled training data; conjoining a classifier to the topic model to create a classifier model; programmatically training the classifier model using labeled training data; receiving target call transcript data; programmatically determining at least one of one or more topics of the target call or one or more classifications of the target call; and digitally storing the target call transcript data with additional data indicating the determined topics and/or classifications of the target call.


