Attribute-Based Visual Search Using Disentangled Encoders

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for interactive image retrieval in online shopping, such as for fashion and furniture, face challenges due to entangled visual attribute representations in the embedding space, leading to uncontrollable and unintended changes in search results when modifying specific attributes, like changing the color of a shirt, which can inadvertently alter other aspects like sleeve type.

Innovation Solution

The development of attribute-specific disentangled encoders that learn separate subspaces for each visual attribute, allowing for controlled manipulation and retrieval of images that maintain overall visual similarity while modifying specific attributes, using techniques like convolutional neural networks (CNNs) and loss functions to disentangle representations and enable attribute-specific operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If visual attribute representations are used in embedding space for image retrieval, then search functionality is enabled, but attribute representations become entangled causing uncontrollable changes in search results

Engineering Contradiction:
Improvesearch functionalityVSAvoidcontrol over search results
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent divides the entangled visual attribute representations into separate, independent attribute-specific subspaces. Each visual attribute (e.g., color, pattern, style) is represented in its own dedicated subspace, allowing independent manipulation of each attribute without affecting others. This segmentation resolves the entanglement problem while maintaining comprehensive search functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and isolates specific visual attribute representations from the mixed embedding space into separate attribute-specific subspaces. By taking out individual attributes from the entangled representation, the system enables precise control over each attribute during image retrieval, preventing unintended changes to other attributes.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If attribute-specific disentangled encoders are used, then control over individual attributes is improved, but system complexity increases

Engineering Contradiction:
Improvecontrol over search resultsVSAvoidencoder architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent employs a universal encoder architecture that can process multiple visual attributes simultaneously while maintaining their independence through separate subspaces. This multi-functional encoder handles various image retrieval tasks (similarity search, attribute manipulation, complementary content retrieval) using the same disentanglement framework, reducing overall system complexity despite the sophisticated internal structure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11829445B1Attribute-based content selection and search
Publication Date: 2023.11.28 AMAZON TECH INC
  • US11829445B1 patent drawing
  • US11829445B1 patent drawing
  • US11829445B1 patent drawing

AI summary

Systems and techniques are generally described for attribute-based content selection and search. In some examples, a graphical user interface (GUI) may display an image of a first product comprising a plurality of visual attributes. In some further examples, the GUI may display at least a first control button with data identifying a first visual attribute of the plurality of visual attributes. In some cases, a first selection of the first control button may be received. In some examples, a first plurality of products may be determined based at least in part on the first selection of the first control button. The first plurality of products may be determined based on a visual similarity to the first product, and a visual dissimilarity to the first product with respect to the first visual attribute. In some examples, the first plurality of products may be displayed on the GUI.