Pitman-Yor Topic Model Pre-Seeded Keyword Groups

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional admixture models for natural language processing are limited by their inability to utilize labeled training data or prior knowledge of themes, requiring large amounts of unlabelled data and extensive human intervention for interpretation, making them inefficient for classifying phone conversations.

Innovation Solution

A classification model based on a hierarchical Pitman-Yor process that incorporates pre-seeded topics and uses unlabeled data for training, allowing for the development of a predictive and interpretable model capable of leveraging prior knowledge, with a Bayesian Belief Network structure that enables efficient training and interpretation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If conventional admixture models are used for topic modeling, then the model can operate without labeled training data, but the model requires large amounts of unlabeled data and extensive human intervention for interpretation

Engineering Contradiction:
Improveease of model developmentVSAvoidamount of data required
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-seeding the topic model with keyword groups before training. These keyword groups serve as initial topic representations that guide the model learning process, allowing the model to leverage prior knowledge and reduce the amount of unlabeled data needed for effective topic discovery and interpretation.

Inventive Principle:
Principle #10Preliminary action

2Extent of automation

If conventional admixture models are used for topic modeling, then the model can operate without labeled training data, but extensive human intervention is required to interpret output and ascribe meaning to topics

Engineering Contradiction:
Improveautomation of topic discoveryVSAvoidtime for human interpretation
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-seeding the topic model with keyword groups before training. These keyword groups serve as initial topic representations that guide the model learning process, allowing the model to leverage prior knowledge and reduce the amount of unlabeled data needed for effective topic discovery and interpretation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback through active learning loops where the model's topic assignments are evaluated and used to refine keyword groupings. This iterative feedback process automatically improves topic interpretation quality over time, reducing the need for manual human intervention while maintaining high interpretability.

Inventive Principle:
Principle #23Feedback

3Reliability

If a highly predictive model is developed using conventional techniques, then the model can achieve good classification accuracy, but the model lacks interpretability and cannot be audited for fairness or correctness

Engineering Contradiction:
Improvepredictive accuracyVSAvoidmodel interpretability
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by separating the model into distinct interpretable components: topic models that identify thematic elements, keyword groups that represent meaningful concepts, and classification layers that combine these elements. This segmented architecture maintains predictive accuracy while enabling auditability of each component's contribution to final predictions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces topic representations as intermediary elements between raw input data and final classifications. These topics serve as interpretable mediators that bridge the gap between complex predictive processing and human-understandable reasoning, allowing auditors to trace decision pathways through meaningful thematic categories.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If labeled training data is used to improve model accuracy, then the model can achieve better classification performance, but the cost of curating labeled data becomes prohibitively expensive

Engineering Contradiction:
Improveclassification accuracyVSAvoidcost of data curation
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-seeding the topic model with keyword groups before training. These keyword groups serve as initial topic representations that guide the model learning process, allowing the model to leverage prior knowledge and reduce the amount of unlabeled data needed for effective topic discovery and interpretation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback through active learning loops where the model's topic assignments are evaluated and used to refine keyword groupings. This iterative feedback process automatically improves topic interpretation quality over time, reducing the need for manual human intervention while maintaining high interpretability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12141669B2Pitman-yor process topic modeling pre-seeded by keyword groupings
Publication Date: 2024.11.12 INVOCA INC
  • US12141669B2 patent drawing
  • US12141669B2 patent drawing
  • US12141669B2 patent drawing

AI summary

In one embodiment, the disclosed technology involves: digitally generating and storing a machine learning statistical topic model in computer memory, the topic model being programmed to model call transcript data representing words spoken on a call as a function of one or more topics of a set of topics that includes pre-seeded topics and non-pre-seeded topics; programmatically pre-seeding the topic model with a set of keyword groups; programmatically training the topic model using unlabeled training data; conjoining a classifier to the topic model to create a classifier model; programmatically training the classifier model using labeled training data; receiving target call transcript data; programmatically determining at least one of one or more topics of the target call or one or more classifications of the target call; and digitally storing the target call transcript data with additional data indicating the determined topics and/or classifications of the target call.