Substitute Identification Model Using Semantic and Image Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying suitable substitutes for items in retail settings are inefficient, often relying on manual processes that are time-consuming and may not provide quick enough turn-around for notifications, especially in automatic replenishment or price drop scenarios, and can be flawed by including complementary products due to reliance on 'viewed-also-viewed' data alone.

Innovation Solution

A two-stage machine learning model that combines 'viewed-also-viewed', 'bought-also-bought', and 'viewed-ultimately-bought' data with semantic and image similarity features to determine substitutive probabilities, using a logistic regression framework to improve accuracy and exclude complementary items.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review processes are used to identify suitable substitutes, then accuracy of substitute identification can be maintained through visual inspection, but the process becomes time-consuming and cannot provide quick turn-around for notifications

Engineering Contradiction:
Improveaccuracy of substitute identificationVSAvoidturn-around time for notifications
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual visual inspection (mechanical human process) with an automated machine learning system that uses multiple data sources and algorithms to identify substitute items. This substitution enables rapid processing of large datasets while maintaining or improving identification accuracy through systematic analysis of viewed-also-viewed, bought-also-bought, and viewed-ultimately-bought data along with semantic and image features.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated systems use only 'viewed-also-viewed' data to identify substitutes, then processing speed is improved, but accuracy deteriorates due to inclusion of complementary products

Engineering Contradiction:
Improveprocessing speed of substitute identificationVSAvoidaccuracy of substitute identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges multiple data sources and feature types into a unified machine learning model. It combines viewed-also-viewed data with bought-also-bought data and viewed-ultimately-bought data, and integrates semantic features and image similarity features. This combination allows the system to distinguish between complementary products and true substitutes by analyzing patterns across multiple dimensions, thereby improving accuracy while maintaining automated processing speed.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The machine learning model acts as an intermediary that processes and synthesizes information from multiple data sources. Rather than directly using only viewed-also-viewed data, the model mediates between multiple data types and produces a refined substitute identification that filters out complementary products while maintaining processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If comprehensive data analysis is performed to improve substitute identification accuracy, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of substitute identificationVSAvoidcomplexity of machine learning system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex analysis task into distinct components: viewed-also-viewed analysis, bought-also-bought analysis, viewed-ultimately-bought analysis, semantic feature extraction, and image similarity analysis. Each component processes a specific type of data or feature independently, then the results are integrated by the machine learning model. This segmentation makes the overall system more manageable and implementable while achieving high identification accuracy.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10769694B2Systems and methods for identifying candidates for item substitution
Publication Date: 2020.09.08 WALMART APOLLO LLC
  • US10769694B2 patent drawing
  • US10769694B2 patent drawing
  • US10769694B2 patent drawing

AI summary

Systems and methods including one or more processors and one or more non-transitory computer-readable media having computing instructions that are configured to run on the one or more processors and perform acts of receiving a test set comprising potential candidate items for substitution for a target item, determining association scores for each of the potential candidate items in the test set, determining one or more semantic similarity features of the potential candidate items in the test set, determining one or more image similarity features of the potential candidate items in the test set, and creating a substitutive probability model by determining a relative contribution of each of the association scores, the semantic similarity features and the image similarity features to a substitutive probability for the potential candidate items in the test set, with reference to a baseline set of the potential candidate items. Additional embodiments are disclosed herein.