Topic Relatedness Scoring for Automated Document Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for automatically generating topics from large corpora face limitations such as models being too broad or too specific, leading to reduced analytics accuracy, and human-generated topics are biased, expensive, and time-consuming to create and maintain, failing to effectively discover related topics and provide precise search results.
Innovation Solution
The system employs probabilistic modeling, specifically a multi-component extension of latent Dirichlet allocation (MC-LDA), to develop multiple topic models with differing numbers of topics, analyzing co-occurring topic IDs to assign relatedness scores and build a hierarchy of topics, enabling precise matching and linking of documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple topic models are used without determining their relatedness, then topic coverage is improved, but the difficulty of finding and discovering features of interest increases
Solution Approach 1:
The patent introduces topic relatedness scores as an intermediary mechanism to connect multiple topic models. These scores serve as a mediator that enables navigation and discovery across topics from different models, resolving the difficulty of finding features while maintaining comprehensive topic coverage.
Solution Approach 2:
The system creates a universal framework that integrates multiple topic models with different granularities. The topic relatedness scoring mechanism provides a common interface for discovering features across diverse topic models, enabling one system to serve multiple topic exploration needs simultaneously.
2Productivity
If automatically generated topics are used, then productivity is improved, but analytics accuracy decreases due to models being too broad or too specific
Solution Approach 1:
The system dynamically adjusts topic granularity by employing multiple topic models with different numbers of topics. Users can navigate between broad and specific topics based on their needs, allowing the system to adapt its level of detail dynamically rather than being fixed at a single granularity level.
Solution Approach 2:
The patent segments the topic space into multiple models with different granularities. Instead of using a single monolithic topic model, the system divides topics into multiple models ranging from broad to specific, allowing users to select the appropriate level of detail for their analytical needs.
3Measurement precision
If human generated topics are used, then analytics accuracy is improved, but cost and time consumption increase
Solution Approach 1:
The system enables self-service topic generation by automatically creating multiple topic models with different granularities. The automated topic relatedness scoring mechanism allows the system to self-organize and self-describe its topic structure without human intervention, eliminating the time and cost of manual topic creation while maintaining analytical accuracy.
Solution Approach 2:
The system uses topic co-occurrence feedback to automatically determine topic relatedness. By analyzing how topics co-occur across documents, the system automatically learns and adjusts topic relationships, replacing manual topic curation with an automated feedback-driven process that maintains accuracy without human time investment.
4Measurement precision
If human generated topics are used, then analytics accuracy is improved, but device complexity and maintenance cost increase
Solution Approach 1:
The patent replaces the mechanical process of manual topic creation and maintenance with an automated computational system. The topic models and relatedness scoring are generated and maintained automatically through algorithmic processes, substituting human intellectual labor with automated machine learning methods that reduce operational complexity.
Data Source
AI summary
A computer system and method for automated discovery of topic relatedness are disclosed. According to an embodiment, topics within documents from a corpus may be discovered by applying multiple topic identification (ID) models, such as multi-component latent Dirichlet allocation (MC-LDA) or similar methods. Each topic model may differ in a number of topics. Discovered topics may be linked to the associated document. Relatedness between discovered topics may be determined by analyzing co-occurring topic IDs from the different models, assigning topic relatedness scores, where related topics may be used for matching/linking a feature of interest. The disclosed method may have an increased disambiguation precision, and may allow the matching and linking of documents using the discovered relationships.


