Cross-Modality Transformer Pre-training for Fine-Grained Fashion Representation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vision-linguistic models are sub-optimal for fine-grained representation learning in domain-specific tasks like fashion, as they primarily focus on coarse representations, making them unfavorable for attribute-aware tasks such as searching for specific fashion catalog objects.

Innovation Solution

A cross-modality transformer-based model is pre-trained using a multi-modal approach that aligns and masks image patches and text tokens to learn fine-grained representations, employing a 'kaleidoscope' patch strategy and attention-based alignment to bridge the semantic gap between text and image modalities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If general VL models (UNITER, VL-BERT) are used for domain-specific tasks, then the model architecture remains simple and general-purpose, but the representation learning capability deteriorates for fine-grained domain-specific features

Engineering Contradiction:
Improvefine-grained representation learning capabilityVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the image into multiple patches and processes them through separate visual embedding layers, allowing fine-grained regional analysis. The text is segmented into tokens processed through language embedding layers. This segmentation enables the model to capture detailed local features while maintaining overall contextual understanding, resolving the contradiction between fine-grained capability and model complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension by adding domain-specific embedding layers for visual and language modalities separately before fusion. This dimensional expansion allows the model to learn domain-specific representations without fundamentally changing the core transformer architecture, thereby improving adaptability while controlling complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If fine-grained representation learning is implemented through domain-specific embedding layers, then the learning accuracy for domain-specific tasks improves, but the training data requirements and computational resources increase

Engineering Contradiction:
Improveattribute-aware recognition accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary domain-specific pre-training to the embedding layers before fine-tuning on specific tasks. This preliminary action allows the model to learn general domain-specific representations from large-scale data first, then adapt to specific tasks with less data, reducing the overall training data requirement while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different embedding strategies to different parts of the model: domain-specific embedding layers for visual and language inputs, and general transformer layers for processing. This localized application of complexity allows high accuracy in attribute-aware recognition while managing training requirements through selective specialization.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11687835B2Domain specific pre-training of cross modality transformer model
Publication Date: 2023.06.27 INCEPTION AI IP LTD
  • US11687835B2 patent drawing
  • US11687835B2 patent drawing
  • US11687835B2 patent drawing

AI summary

A transformer based vision-linguistic (VL) model and training technique uses a number of different image patches covering the same portion of an image, along with a text description of the image to train the model. The model and pre-training techniques may be used in domain specific training of the model. The model can be used for fine-grained image-text tasks in the fashion domain.