Multimodal Learning with Masking Objectives
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multimodal machine learning models face challenges in generalizing across various input modalities, especially when structured data exhibits non-stationary behavior and modalities are missing, leading to difficulties in building accurate and generalizable models.
Innovation Solution
A multimodal processing system that trains models using modality-specific masking objectives and joint modality similarity-based masking objectives to generate consistent outputs even with partially or completely missing data, reinforcing cross-modal relationships and improving model accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multimodal machine learning models are trained to handle multiple data modalities, then the model's ability to process diverse input types is improved, but the difficulty in building generalizable models increases due to non-stationary behavior in structured data and missing modalities
Solution Approach 1:
The patent segments the training process into two distinct phases: pretraining with masking objectives to handle missing modalities, and fine-tuning for specific tasks. This segmentation allows the model to first learn robust representations that are invariant to missing data, then adapt to specific tasks, thereby improving generalizability while maintaining versatility.
Solution Approach 2:
The patent applies preliminary action by performing pretraining with masking objectives before task-specific fine-tuning. During pretraining, the model learns to generate consistent representations even when modalities are missing, which prepares the model to handle real-world data variations before encountering specific tasks, thus improving reliability.
2Measurement precision
If modality-specific masking objectives are used during pretraining, then the model's accuracy with missing data is improved, but the training complexity increases
Solution Approach 1:
The patent implements a unified masking objective that serves multiple functions simultaneously: it handles missing modalities, learns cross-modal relationships, and generates consistent representations. This multi-functional approach improves accuracy with missing data while avoiding the need for separate training mechanisms for each function, thereby managing training complexity.
Solution Approach 2:
The masking objective incorporates feedback by comparing representations from masked and unmasked inputs and minimizing the difference. This feedback mechanism guides the model to learn robust representations that are invariant to missing data, improving accuracy while using a straightforward optimization approach that manages training complexity.
3Measurement precision
If joint modality similarity-based masking objectives are applied, then cross-modal relationships are reinforced and model accuracy increases, but computational expenses increase
Solution Approach 1:
The patent merges the learning of cross-modal relationships into the pretraining phase through joint modality similarity objectives, rather than requiring separate post-processing or additional training stages. This combining approach reinforces cross-modal relationships and improves accuracy while consolidating computational work into the existing pretraining pipeline, managing overall computational expenses.
4Ease of operation
If structured data is converted to unstructured formats for processing, then uniform processing is simplified, but computational expenses increase and model accuracy may decrease
Solution Approach 1:
The patent segments the data processing approach by maintaining separate handling pathways for structured and unstructured data through modality-specific encoders, rather than converting everything to a single format. This segmentation allows each data type to be processed in its optimal format, avoiding the computational overhead and potential information loss of conversion while still enabling unified model processing through modality-specific masking objectives.
Data Source
AI summary
Aspects of the disclosure are directed to a multimodal processing system for processing both structured and un-structured data. Real-world data is not always consistent in form or content. The multimodal processing system includes model that can be trained to account for this characteristic of real-world data, by selectively masking data of different modalities during pretraining to learn outputs that are the same or comparable between the masked and un-masked inputs. The model is trained according to modality-specific masking objectives computed for each modality of data and joint modality similarity-based masking objectives for a joint representation of the data across all modalities. The system provides consistent and accurate input, even when input data may have substantial portions of data from different modalities missing. Cross-modal relationships in data are reinforced by the model as different portions of data are masked, contributing to an overall increase in model accuracy versus other approaches.


