Automated Product Categorization via ML Text Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual product categorization in e-commerce is time-consuming and costly, leading to inconsistencies and incompatibilities across different categorization systems, making it difficult for merchants to keep product classifications relevant and up-to-date.
Innovation Solution
An automated system that uses text metadata fields such as product titles, descriptions, and brand names to categorize products by estimating probabilities through machine learning classifiers, allowing for faster and more accurate categorization and association with relevant categories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual categorization is used, then product categorization can be performed with simple systems, but it is time-consuming and costly
Solution Approach 1:
The system enables self-service automated categorization by allowing the classification system to automatically categorize products using machine learning algorithms and text metadata analysis without requiring manual human intervention for each product categorization task
Solution Approach 2:
The patent replaces the mechanical manual categorization process with an automated electronic system that uses machine learning classifiers, text metadata analysis, and probability estimation algorithms to perform categorization tasks that were previously done manually
2Measurement precision
If manual categorization is used, then categorization can be performed with simple systems, but accuracy and consistency across different categorization systems deteriorate
Solution Approach 1:
The patent introduces intermediary components including text metadata fields as intermediaries between products and categories, and machine learning classifiers as intermediaries that process the metadata and estimate probabilities to bridge the gap between product information and category assignments
Solution Approach 2:
The system changes parameters by using multiple text metadata fields (title, description, brand name) instead of single-field categorization, and by implementing probability-based category selection with threshold parameters to improve categorization accuracy and consistency
3Productivity
If automated categorization is implemented, then productivity and accuracy improve, but system complexity increases
Solution Approach 1:
The patent segments the categorization process into distinct components: text metadata extraction from products, feature vector generation from metadata, machine learning classification processing, and category assignment based on probability thresholds. This segmentation allows each component to be optimized independently while maintaining overall system productivity
Data Source
AI summary
A system and method is described for large-scale, automated classification of products. The system and method receives information about products, wherein such information includes one or more text metadata fields associated with each product, receives a set of categories, and automatically selects one or more categories from the set of categories to which each product belongs based upon at least one of the one or more text metadata fields associated with each product. A machine learning classifier may be used to automatically select the one or more categories to which each product belongs by operating upon a feature vector for each product derived from text metadata fields of the product description. The machine learning classifier may be trained using a set of pre-categorized product descriptions. The product-category associations generated by the system and method can be used to improve search engine results or product recommendations to consumers.


