Automated Corpus Theme Detection With Diverse Phrase Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying themes within large data sets are inefficient, lacking diversity, and require significant time and resources, often leading to late or irrelevant evaluations of user-provided feedback.

Innovation Solution

Employ unsupervised machine learning and natural language processing techniques to cluster and rank candidate phrases within a corpus of information, using centroid-based clustering and diversity-based ranking to identify themes from user submissions, enabling automated theme detection across various services and content types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional manual methods are used to identify themes within large data sets, then evaluation thoroughness may be maintained, but time consumption and resource requirements increase significantly

Engineering Contradiction:
Improvetheme identification accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical analysis with automated machine learning systems. The system uses unsupervised learning algorithms to automatically identify themes, extract candidate phrases, and rank them by relevance, eliminating the need for manual data analysis while maintaining or improving accuracy through consistent algorithmic application across the entire data set.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service theme identification by automatically processing data sets without human intervention. The machine learning model autonomously performs clustering, phrase extraction, and ranking operations, allowing the system to serve its own analytical needs and scale independently of human resources.

Inventive Principle:
Principle #25Self-service

2Loss of information

If comprehensive analysis of all data is performed, then complete theme coverage is achieved, but processing resources and time requirements increase

Engineering Contradiction:
Improvetheme coverage completenessVSAvoidprocessing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent segments the data analysis process into distinct stages: initial clustering of data points, extraction of candidate phrases from cluster centroids, and sequential ranking of themes. This segmentation allows the system to process large data sets efficiently by breaking down the comprehensive analysis task into manageable, automated steps that maintain completeness while improving productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial analysis by initially focusing on cluster centroids and representative phrases, then progressively refining theme identification through ranking. This approach achieves sufficient theme coverage without requiring exhaustive analysis of every single data point, thereby improving processing efficiency while maintaining adequate information completeness.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If diverse theme identification methods are used, then theme diversity improves, but system complexity increases

Engineering Contradiction:
Improvetheme detection diversityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal machine learning framework that handles multiple theme detection functions through a single system. The same unsupervised learning model performs clustering, phrase extraction, and ranking operations, providing diverse theme identification capabilities without requiring separate specialized systems for each function, thus maintaining versatility while controlling complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If automated theme detection is implemented, then processing speed improves, but resource requirements increase

Engineering Contradiction:
Improvetheme identification speedVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs partial processing by focusing computational resources on analyzing cluster centroids and representative phrases rather than every individual data point. This approach achieves fast automated theme detection while reducing computational resource consumption by processing only the most informative portions of the data set.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12423345B2Theme detection within a corpus of information
Publication Date: 2025.09.23 AMAZON TECH INC
  • US12423345B2 patent drawing
  • US12423345B2 patent drawing
  • US12423345B2 patent drawing

AI summary

Systems and methods are used to detect underlying themes from a collection of documents at an aggregated level. A representative set of documents may be selected from a cluster of documents, with the representative set of documents corresponding to a general theme of the cluster. Candidate theme phrases may then be extracted from the documents and used to generate document embeddings and candidate phrase embeddings, which may be ranked, such as with a diversity-based ranking approach. Certain candidates may be selected from the ranking. Each of the documents forming the representative set may then be concatenated and a query embedding may be generated and ranked against the candidate phrases. In this manner, a collection of phrases associated with both the general underlying theme of the cluster, along with granular topics associated with that theme, may be identified.