Masked Distillation for Accurate Multilingual Vision-Language Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing vision-language models suffer from inaccuracies in generating vision-language outputs, particularly for underrepresented languages, and are computationally inefficient due to reliance on noisy alt-text data and extensive training on large datasets.
Innovation Solution
A masked distillation system that combines masked distillation and contrastive image-text pretraining to train a vision-language model, leveraging features from foundation models to improve accuracy and efficiency by learning linear projections from shared embedding spaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing vision-language models are trained using noisy alt-text data and extensive training on large datasets, then the model can perform basic vision-language functions, but the quality of vision-language outputs deteriorates and computational efficiency worsens
Solution Approach 1:
The patent applies preliminary action by pre-training foundation models (vision encoder and language model) separately on large-scale datasets before distilling knowledge into the vision-language model. This pre-training phase prepares the foundation models with robust visual and linguistic capabilities, which are then transferred to the VL model through distillation, improving output quality while reducing the need for extensive noisy data training
Solution Approach 2:
The patent uses an intermediary approach by introducing foundation models as intermediate components that mediate between raw data and the final vision-language model. The vision encoder and language model serve as intermediaries that process and refine information before integration, enabling high-quality outputs without direct training on noisy alt-text data
2Ease of manufacture
If vision-language models are trained with extensive datasets and noisy alt-text data, then basic functionality is achieved, but manufacturing precision of vision-language outputs deteriorates
Solution Approach 1:
The patent applies the taking out principle by extracting and removing the problematic training step that uses noisy alt-text data. Instead of training the vision-language model directly on noisy data, the method extracts knowledge from pre-trained foundation models through distillation, thereby eliminating the source of inaccuracy while maintaining training simplicity
Solution Approach 2:
The patent uses copying by replicating the successful training approach of foundation models and applying it to the vision-language model through knowledge distillation. The VL model copies the learned representations and processing mechanisms from the foundation models, achieving high precision outputs without directly processing noisy training data
3Measurement precision
If masked distillation and contrastive pretraining are used to train the vision-language model, then accuracy of vision-language outputs improves, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by performing contrastive pretraining of the vision encoder and language model separately before the distillation phase. This pretraining establishes accurate visual and linguistic representations in advance, allowing the subsequent distillation process to focus on integration rather than fundamental learning, thereby achieving high accuracy with manageable complexity
Solution Approach 2:
The patent uses segmentation by dividing the training process into distinct phases: contrastive pretraining of the vision encoder, pretraining of the language model, and final distillation to integrate them. This segmentation allows each component to be optimized independently with appropriate techniques, achieving high accuracy while managing overall system complexity through modular development
Data Source
AI summary
The present disclosure relates to systems, non-transitory computer-readable media, and methods for training and implementing a vision-language model using masked distillation and contrastive image-text training. In particular, in one or more embodiments, the disclosed systems generate, utilizing a vision encoder, an image embedding from a masked digital image comprising a digital image with one or more masked patches. In some embodiments, the disclosed systems generate, utilizing a text encoder, a text embedding from a masked text phrase. In one or more embodiments, the disclosed systems generate, utilizing the vision-language model from the image embedding and the text embedding, a predicted text reconstruction of the text description and a predicted image reconstruction of the digital image. In some embodiments, the disclosed systems modify parameters of the vision-language model according to a masked distillation loss between the predicted text reconstruction and a text reconstruction generated by a pretrained large language model.


