Taxonomy Topic Determination via Query Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The creation and modification of taxonomies using traditional machine learning techniques are laborious, time-consuming, and expensive, and face challenges in classifying nuanced and non-compositional queries due to linguistic complexities such as polysemy and synonymy, leading to errors and overgeneralizations.

Innovation Solution

The method determines topics by generating two layers of associations: tokens are associated and grounded in a corpus of queries, and then topics are associated with nodes of a taxonomy, allowing for efficient and simple association of nodes with topics, enabling quick evaluation and modification of taxonomies without the need for extensive training data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional machine learning techniques are used to create and modify taxonomies, then the taxonomy can be built with established methods, but the process becomes laborious, time-consuming, and expensive

Engineering Contradiction:
Improvetaxonomy accuracyVSAvoidtaxonomy creation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent introduces an intermediary process that uses query analysis and topic modeling as intermediate steps between raw data and taxonomy creation. By analyzing user queries to extract topics and associations, the system creates a bridge that automates the taxonomy development process, reducing manual labor while maintaining accuracy through data-driven topic identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical, manual process of taxonomy creation with an automated computational system. Instead of manually curating and organizing taxonomy nodes, the system uses natural language processing, query analysis, and machine learning algorithms to automatically generate and update taxonomies, significantly reducing time and labor costs while maintaining or improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If traditional machine learning techniques are used for text-based classifications, then the taxonomy can be populated with classified content, but linguistic complexities such as polysemy and synonymy lead to difficulties in classification

Engineering Contradiction:
Improveclassification throughputVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the system continuously analyzes user queries and interactions with the taxonomy. This feedback loop allows the system to learn from actual usage patterns, refine topic models, and improve classification accuracy over time. The system adapts to linguistic complexities by incorporating real-world query data that reflects how users actually express polysemous and synonymous concepts.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameters used for classification from traditional keyword matching to topic-based associations derived from query analysis. By transforming the classification approach to use semantic topics and associations extracted from user queries, the system better handles polysemy and synonymy, improving classification accuracy while maintaining throughput through automated processing.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If image-based classifications are used to populate taxonomies, then visual content can be categorized, but errors and overgeneralizations occur based on how the models are trained

Engineering Contradiction:
Improvecontent type coverageVSAvoidclassification accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal taxonomy system that can handle multiple content types including text, images, and queries through a unified topic-based approach. The system uses topic models that can be applied across different modalities, allowing the same taxonomy structure to serve diverse content types while maintaining accuracy through consistent topic-based classification rather than modality-specific approaches.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If extensive training data collection is performed to improve taxonomy accuracy, then the taxonomy can be trained more precisely, but the process becomes more expensive and time-consuming

Engineering Contradiction:
Improvetaxonomy training accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent enables the taxonomy system to self-improve by automatically analyzing user queries and interactions to generate training data and refine topic models. Instead of requiring extensive external data collection, the system leverages its own operational data from user queries and taxonomy usage to continuously learn and improve accuracy, reducing the need for additional training data while maintaining or enhancing precision.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240320254A1Determining topics for taxonomies
Publication Date: 2024.09.26 PINTEREST INC
  • US20240320254A1 patent drawing
  • US20240320254A1 patent drawing
  • US20240320254A1 patent drawing

AI summary

Systems and methods for determining one or more topics that may be associated and/or grounded with a node of a taxonomy to facilitate the creation and/or modification of the taxonomy. The one or more topics can be determined by generating two layers of associations and/or groundings. In a first layer, tokens can be associated and/or grounded in a corpus of queries, and in a second layer, topics can be associated and/or grounded in the tokens. The topics can then be associated with and/or grounded in nodes of a taxonomy which can facilitate access to content items stored and maintained by an online service. Further, in exemplary implementations where the content items include associations and/or mappings to the corpus of queries, the nodes of the taxonomy (which are associated with one or more topics) can be transitively mapped to the content items.