Conversational Shopping System Using Text and Image Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing conversational shopping systems in e-commerce are not natural, relying on complicated category-specific jargons and failing to allow customers to visualize furniture or multiple product categories together, limiting their ability to converse effectively and discover desired products.
Innovation Solution
A system using machine learning to process both text and image data, allowing customers to submit reference images and queries, which computes text and image embeddings to determine target products, enabling natural language-based shopping and visualization of multiple product categories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If existing conversational shopping systems use rule-based flows and hardcoded filters, then item retrieval can be performed, but the system becomes unnatural and requires complicated category-specific jargons that customers are not familiar with
Solution Approach 1:
The patent replaces rule-based mechanical systems with machine learning models that process natural language and image embeddings. The system uses trained models to understand customer intent from natural language queries and reference images, eliminating the need for hardcoded filters and category-specific jargons while maintaining retrieval functionality.
Solution Approach 2:
The system transforms the approach by changing from discrete filter parameters to continuous embedding space representations. Customer queries and product attributes are converted into vector embeddings, allowing for more flexible and natural parameter matching that doesn't require predefined categorical structures.
2Adaptability or versatility
If existing systems use textual search with hardcoded filters, then item retrieval is possible, but customers cannot visualize furniture or multiple product categories together
Solution Approach 1:
The patent merges text and image modalities into a unified retrieval framework. Both the customer's reference image and textual query are processed together through the machine learning model, allowing the system to understand both visual appearance and descriptive context simultaneously for more accurate product recommendations.
Solution Approach 2:
The system adds an image dimension to traditional text-based search by incorporating image embeddings. This allows customers to visualize products and their contexts (such as furniture in a room setting) while maintaining the textual query capability, creating a multi-dimensional retrieval approach.
3Adaptability or versatility
If voice-based search uses rule-based flows, then conversation can be emulated, but natural conversation and building furnished rooms is not allowed
Solution Approach 1:
The patent replaces rule-based voice search mechanisms with machine learning models that process both audio input and image data. The system uses trained models to understand conversational intent and visual references simultaneously, enabling natural conversations about furnishing rooms without requiring predefined dialogue flows.
4Measurement precision
If existing systems use category-specific jargons, then precise item retrieval can be achieved, but customers are not familiar with these terms making the system difficult to use
Solution Approach 1:
The system introduces machine learning models as intermediaries between customer natural language and product database structures. The models translate colloquial customer descriptions into meaningful product attributes without requiring customers to know specific jargons, while still achieving accurate retrieval by mapping to the underlying product data structure.
Data Source
AI summary
Systems and methods conversational shopping based on machine learning using both text and image data are disclosed. In some embodiments, a disclosed method includes: obtaining, from a computing device, a search request identifying a reference image representing a first product and a textual query associated with the reference image; computing a text embedding in an embedding space based on the textual query; computing an image embedding in the embedding space based on the reference image; determining, based on at least one machine learning model, a target image representing a second product based on the text embedding and the image embedding; and transmitting, to the computing device, the target image in response to the search request.


