Automated Product Categorization via ML Text Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual product categorization in e-commerce is time-consuming and costly, leading to inconsistencies and incompatibilities across different categorization systems, making it difficult for merchants to keep product classifications relevant and up-to-date.

Innovation Solution

An automated system that uses text metadata fields such as product titles, descriptions, and brand names to categorize products by estimating probabilities through machine learning classifiers, allowing for faster and more accurate categorization and association with relevant categories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual categorization is used, then product categorization can be performed with simple systems, but it is time-consuming and costly

Engineering Contradiction:
Improvecategorization speedVSAvoidtime for manual categorization
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables self-service automated categorization by allowing the classification system to automatically categorize products using machine learning algorithms and text metadata analysis without requiring manual human intervention for each product categorization task

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual categorization process with an automated electronic system that uses machine learning classifiers, text metadata analysis, and probability estimation algorithms to perform categorization tasks that were previously done manually

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual categorization is used, then categorization can be performed with simple systems, but accuracy and consistency across different categorization systems deteriorate

Engineering Contradiction:
Improvecategorization accuracyVSAvoidcomplexity of automated classification system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces intermediary components including text metadata fields as intermediaries between products and categories, and machine learning classifiers as intermediaries that process the metadata and estimate probabilities to bridge the gap between product information and category assignments

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes parameters by using multiple text metadata fields (title, description, brand name) instead of single-field categorization, and by implementing probability-based category selection with threshold parameters to improve categorization accuracy and consistency

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated categorization is implemented, then productivity and accuracy improve, but system complexity increases

Engineering Contradiction:
Improvecategorization efficiencyVSAvoidcomplexity of machine learning system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the categorization process into distinct components: text metadata extraction from products, feature vector generation from metadata, machine learning classification processing, and category assignment based on probability thresholds. This segmentation allows each component to be optimized independently while maintaining overall system productivity

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10528907B2Automated categorization of products in a merchant catalog
Publication Date: 2020.01.07 YAHOO AD TECH LLC
  • US10528907B2 patent drawing
  • US10528907B2 patent drawing
  • US10528907B2 patent drawing

AI summary

A system and method is described for large-scale, automated classification of products. The system and method receives information about products, wherein such information includes one or more text metadata fields associated with each product, receives a set of categories, and automatically selects one or more categories from the set of categories to which each product belongs based upon at least one of the one or more text metadata fields associated with each product. A machine learning classifier may be used to automatically select the one or more categories to which each product belongs by operating upon a feature vector for each product derived from text metadata fields of the product description. The machine learning classifier may be trained using a set of pre-categorized product descriptions. The product-category associations generated by the system and method can be used to improve search engine results or product recommendations to consumers.