Multi-modal Product Embedding Generator for Recommendation Services

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing recommendation systems rely on single-modal features, limiting their ability to generate product embeddings that are compatible with multiple types of product information, such as images and text, which restricts their applicability across different recommendation services.

Innovation Solution

A multi-task trained machine learning model is developed to generate product embeddings from multiple types of product information, including images and text, enabling compatibility with various query types and optimizing embeddings for multiple engagement types, thus supporting multiple recommendation services.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing systems use single-modal features for generating product embeddings, then the system complexity is reduced, but the compatibility with multiple types of product information (images, text, etc.) is limited

Engineering Contradiction:
Improvecompatibility with multiple types of product informationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a unified embedding generation system that processes multiple modalities (images, text, and other product information) through a single multi-modal product embedding generator. This universal system replaces multiple separate single-modal systems, enabling compatibility with various recommendation services that accept different query types while managing complexity through integrated architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple separate systems are used to support different recommendation services, then the adaptability to different query types is improved, but the infrastructure and maintenance costs increase

Engineering Contradiction:
Improvesupport for multiple recommendation servicesVSAvoidinfrastructure and maintenance costs
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges multiple separate recommendation service systems into a single multi-modal product embedding generator that serves all services. By combining image processing, text processing, and other information processing into one unified system, the patent reduces infrastructure requirements and maintenance costs while maintaining the ability to support diverse recommendation services through multi-task training.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If single-modal embeddings are used, then the manufacturing precision of embedding generation is simplified, but the accuracy of product recommendations across diverse inputs is reduced

Engineering Contradiction:
Improveaccuracy of product recommendationsVSAvoidembedding generation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates composite product embeddings by integrating features from multiple modalities (images, text, and other product information) into a unified embedding representation. This composite approach combines the strengths of different feature types to improve recommendation accuracy across diverse query inputs, with the multi-modal generator synthesizing heterogeneous data into coherent embeddings through learned feature fusion.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20230252550A1Multi-modal product embedding generator
Publication Date: 2023.08.10 PINTEREST INC
  • US20230252550A1 patent drawing
  • US20230252550A1 patent drawing
  • US20230252550A1 patent drawing

AI summary

Described are systems and methods for providing a multi-tasked trained machine learning model that may be configured to generate product embeddings from multiple types of product information. The exemplary product embeddings may be generated for a corpus of products (e.g., products included in a product catalog, etc.) based on both image information and text information associated with each respective product. Accordingly, the generated product embeddings may be compatible with learned representations of the different types of product information (e.g., image information, text information, etc.) and may be used to create a product index, which can be used to determine and serve product recommendations in connection with multiple different recommendation services that may be configured to receive different types of inputs (e.g., a single image, multiple images, text-based information, etc.).