Product Similarity Identification via Topic Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying product substitutes in retail are inconsistent and unreliable, as they rely on expert judgment or statistical models that require historical data and specific conditions, often failing to account for consumer behavior randomness and lack of data.
Innovation Solution
A computer-implemented method that queries internet databases, performs topic modeling on product descriptions using techniques like latent semantic analysis, nonnegative matrix factorization, and singular value decomposition to identify sets of similar products based on their descriptive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If expert judgment approach is used to identify substitute products, then business expertise and knowledge can be leveraged, but inconsistencies in knowledge and skill across the company arise and information maintenance becomes difficult
Solution Approach 1:
The patent replaces the manual expert judgment approach with an automated statistical model that objectively analyzes historical sales data to identify substitute products. This eliminates human inconsistency and reduces the operational burden of maintaining substitution information, as the system automatically updates based on new sales data.
Solution Approach 2:
The system enables self-service by automatically gathering and analyzing sales data to identify product substitutions without requiring continuous expert intervention. The model autonomously maintains and updates substitution relationships based on evolving sales patterns.
2Reliability
If statistical approach is used to determine product substitutes, then objective analysis of sales data is achieved, but the model requires specific historical conditions and data availability that are often not met
Solution Approach 1:
The patent applies partial action by using available sales data even when it doesn't meet ideal statistical conditions. The model can identify substitutions with limited data by focusing on the most relevant patterns, rather than requiring complete fulfillment of all statistical assumptions.
Solution Approach 2:
The system adapts to different data conditions by adjusting model parameters and assumptions based on data availability. When historical data is limited or conditions aren't met, the model modifies its approach to still identify meaningful substitution relationships.
3Measurement precision
If correlation model is used to identify substitutes, then pricing and promotional variations can be analyzed, but consumer behavior randomness and non-modeled phenomena reduce measurement confidence
Solution Approach 1:
The patent incorporates feedback mechanisms that continuously monitor and evaluate the confidence levels of substitution identifications. The system uses this feedback to adjust its measurements and declarations, maintaining reliability by only declaring substitutions with sufficient confidence while continuously improving measurement precision through learned patterns.
Data Source
AI summary
Embodiments of the present invention relate to systems and methods for determining sets of products which are similar to each other in terms of consumers' wants and needs. Queries are performed on a particular product. Documents relating to the query are received and stored. A dictionary is created from the received documents, whereby the documents, which are text files, are scrubbed of certain data to create a scrubbed text file. Topic modeling is then performed on the cleansed text file. Various methods can be used to perform topic modeling, including, but not limited to, latent semantic analysis, nonnegative matrix factorization, and singular value decomposition.


