Masked Distillation for Accurate Multilingual Vision-Language Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vision-language models suffer from inaccuracies in generating vision-language outputs, particularly for underrepresented languages, and are computationally inefficient due to reliance on noisy alt-text data and extensive training on large datasets.

Innovation Solution

A masked distillation system that combines masked distillation and contrastive image-text pretraining to train a vision-language model, leveraging features from foundation models to improve accuracy and efficiency by learning linear projections from shared embedding spaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing vision-language models are trained using noisy alt-text data and extensive training on large datasets, then the model can perform basic vision-language functions, but the quality of vision-language outputs deteriorates and computational efficiency worsens

Engineering Contradiction:
Improvequality of vision-language outputsVSAvoidcomputational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-training foundation models (vision encoder and language model) separately on large-scale datasets before distilling knowledge into the vision-language model. This pre-training phase prepares the foundation models with robust visual and linguistic capabilities, which are then transferred to the VL model through distillation, improving output quality while reducing the need for extensive noisy data training

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses an intermediary approach by introducing foundation models as intermediate components that mediate between raw data and the final vision-language model. The vision encoder and language model serve as intermediaries that process and refine information before integration, enabling high-quality outputs without direct training on noisy alt-text data

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If vision-language models are trained with extensive datasets and noisy alt-text data, then basic functionality is achieved, but manufacturing precision of vision-language outputs deteriorates

Engineering Contradiction:
Improvetraining processVSAvoidaccuracy of vision-language outputs
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent applies the taking out principle by extracting and removing the problematic training step that uses noisy alt-text data. Instead of training the vision-language model directly on noisy data, the method extracts knowledge from pre-trained foundation models through distillation, thereby eliminating the source of inaccuracy while maintaining training simplicity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses copying by replicating the successful training approach of foundation models and applying it to the vision-language model through knowledge distillation. The VL model copies the learned representations and processing mechanisms from the foundation models, achieving high precision outputs without directly processing noisy training data

Inventive Principle:
Principle #26Copying

3Measurement precision

If masked distillation and contrastive pretraining are used to train the vision-language model, then accuracy of vision-language outputs improves, but device complexity increases

Engineering Contradiction:
Improveaccuracy of vision-language outputsVSAvoidmodel architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing contrastive pretraining of the vision encoder and language model separately before the distillation phase. This pretraining establishes accurate visual and linguistic representations in advance, allowing the subsequent distillation process to focus on integration rather than fundamental learning, thereby achieving high accuracy with manageable complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses segmentation by dividing the training process into distinct phases: contrastive pretraining of the vision encoder, pretraining of the language model, and final distillation to integrate them. This segmentation allows each component to be optimized independently with appropriate techniques, achieving high accuracy while managing overall system complexity through modular development

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250265831A1Building vision-language models using masked distillation from foundation models
Publication Date: 2025.08.21 ADOBE INC
  • US20250265831A1 patent drawing
  • US20250265831A1 patent drawing
  • US20250265831A1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods for training and implementing a vision-language model using masked distillation and contrastive image-text training. In particular, in one or more embodiments, the disclosed systems generate, utilizing a vision encoder, an image embedding from a masked digital image comprising a digital image with one or more masked patches. In some embodiments, the disclosed systems generate, utilizing a text encoder, a text embedding from a masked text phrase. In one or more embodiments, the disclosed systems generate, utilizing the vision-language model from the image embedding and the text embedding, a predicted text reconstruction of the text description and a predicted image reconstruction of the digital image. In some embodiments, the disclosed systems modify parameters of the vision-language model according to a masked distillation loss between the predicted text reconstruction and a text reconstruction generated by a pretrained large language model.