Machine Learning Missing-Token Metadata Augmentation for Product Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Marketplaces face challenges in helping customers find desired products due to the large number of available items, leading to many irrelevant matches in free-form text searches.

Innovation Solution

A machine learning model is trained to determine missing tokens in product metadata using user engagement information, allowing for improved search retrieval by augmenting product metadata with relevant keywords.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If free-form text search is used to search through millions of products, then search coverage is improved, but search accuracy deteriorates due to many irrelevant matches

Engineering Contradiction:
Improvesearch coverageVSAvoidsearch accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces traditional mechanical text matching systems with a machine learning model that uses natural language processing to understand search intent and generate relevant keywords. This substitution enables the system to maintain broad search coverage while significantly improving search accuracy by moving from simple keyword matching to semantic understanding.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the parameters of the search system by dynamically generating and weighting multiple keywords for each product based on user engagement data. Instead of relying on a single static search term, the system adjusts keyword parameters (relevance scores, weights) to optimize both coverage and precision in search results.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If product metadata is augmented with more keywords, then search retrieval accuracy is improved, but data processing complexity increases

Engineering Contradiction:
Improvesearch retrieval accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a self-service system where the machine learning model automatically generates keywords and updates product metadata without requiring manual intervention. The system uses user engagement information to autonomously improve product descriptions, thereby reducing the complexity of data processing while maintaining high search retrieval accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary action by pre-generating keywords and augmenting product metadata before users perform searches. This advance preparation reduces the complexity of real-time search processing, as the system has already organized and optimized product information in advance based on historical user engagement patterns.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If user engagement information is collected and processed to train the model, then search personalization is improved, but data privacy concerns increase

Engineering Contradiction:
Improvesearch personalizationVSAvoiddata privacy concerns
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary engagement information (clicks, views, purchases) required for training the machine learning model while excluding personally identifiable information. This selective extraction allows the system to achieve search personalization while minimizing data privacy concerns by collecting only the minimum necessary data.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250245713A1Systems and methods for training a machine learning model to determine missing tokens to augment product metadata for search retrieval
Publication Date: 2025.07.31 WALMART APOLLO LLC
  • US20250245713A1 patent drawing
  • US20250245713A1 patent drawing
  • US20250245713A1 patent drawing

AI summary

Systems and methods including one or more processors and one or more non-transitory storage devices storing computing instructions configured to run on the one or more processors and perform: receiving user engagement information from a plurality of users, the user engagement information including pairing information; filtering the pairing information based on filtering criteria; training a machine learning model based on the pairing information to determine missing tokens; and modifying metadata for one or more products in a product catalog based on the missing tokens. Other embodiments are disclosed herein.