Multi-modal Product Embedding Generator for Recommendation Services
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing recommendation systems rely on single-modal features, limiting their ability to generate product embeddings that are compatible with multiple types of product information, such as images and text, which restricts their applicability across different recommendation services.
Innovation Solution
A multi-task trained machine learning model is developed to generate product embeddings from multiple types of product information, including images and text, enabling compatibility with various query types and optimizing embeddings for multiple engagement types, thus supporting multiple recommendation services.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing systems use single-modal features for generating product embeddings, then the system complexity is reduced, but the compatibility with multiple types of product information (images, text, etc.) is limited
Solution Approach 1:
The patent implements a unified embedding generation system that processes multiple modalities (images, text, and other product information) through a single multi-modal product embedding generator. This universal system replaces multiple separate single-modal systems, enabling compatibility with various recommendation services that accept different query types while managing complexity through integrated architecture.
2Adaptability or versatility
If multiple separate systems are used to support different recommendation services, then the adaptability to different query types is improved, but the infrastructure and maintenance costs increase
Solution Approach 1:
The patent merges multiple separate recommendation service systems into a single multi-modal product embedding generator that serves all services. By combining image processing, text processing, and other information processing into one unified system, the patent reduces infrastructure requirements and maintenance costs while maintaining the ability to support diverse recommendation services through multi-task training.
3Measurement precision
If single-modal embeddings are used, then the manufacturing precision of embedding generation is simplified, but the accuracy of product recommendations across diverse inputs is reduced
Solution Approach 1:
The patent creates composite product embeddings by integrating features from multiple modalities (images, text, and other product information) into a unified embedding representation. This composite approach combines the strengths of different feature types to improve recommendation accuracy across diverse query inputs, with the multi-modal generator synthesizing heterogeneous data into coherent embeddings through learned feature fusion.
Data Source
AI summary
Described are systems and methods for providing a multi-tasked trained machine learning model that may be configured to generate product embeddings from multiple types of product information. The exemplary product embeddings may be generated for a corpus of products (e.g., products included in a product catalog, etc.) based on both image information and text information associated with each respective product. Accordingly, the generated product embeddings may be compatible with learned representations of the different types of product information (e.g., image information, text information, etc.) and may be used to create a product index, which can be used to determine and serve product recommendations in connection with multiple different recommendation services that may be configured to receive different types of inputs (e.g., a single image, multiple images, text-based information, etc.).


