Multi-Modal Data Labeling Through Inter- and Intra-Modal Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data labeling methods are labor-intensive and inefficient for multi-modal data, limiting the rapid development of machine learning models due to insufficient training data and manual labeling requirements.
Innovation Solution
A computer-implemented method that combines inter-modal and intra-modal label transformations to automate data labeling, using a kernel function to estimate and combine labels for multi-modal data, reducing training time and improving model accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data labeling is used for multi-modal data, then labeling accuracy can be maintained, but labeling efficiency and speed deteriorate significantly
Solution Approach 1:
The system enables semi-supervised learning where the model automatically labels unlabeled data using its own learned representations. The cross-modal contrastive learning framework allows the system to self-improve by leveraging both labeled and unlabeled data, reducing dependency on manual labeling while maintaining accuracy through iterative refinement of label predictions
Solution Approach 2:
The patent introduces cross-modal contrastive learning as an intermediary mechanism that bridges different data modalities. By learning shared representations across modalities, the system can transfer labeling knowledge between modalities, enabling accurate automatic labeling without direct manual intervention for each modality
2Reliability
If more training data is collected for machine learning models, then model accuracy improves, but data labeling time and resources increase
Solution Approach 1:
The system performs preliminary contrastive learning on available labeled data to establish cross-modal relationships before tackling the full labeling task. This preliminary representation learning enables the model to quickly generalize to unlabeled data, reducing the time needed for comprehensive labeling while improving model accuracy through pre-established feature alignments
Solution Approach 2:
The patent employs semi-supervised learning where only a portion of the data requires manual labeling. The model leverages this partial labeled set to learn cross-modal relationships, then applies these relationships to automatically label the remaining excessive unlabeled data, achieving high model accuracy without proportionally increasing manual labeling time
Data Source
AI summary
One or more computer processors extract respective features for each inter-modal sample in an inter-modal dataset, for each intra-modal sample in an intra-modal dataset, and a subsequent sample, wherein the inter-modal dataset and the intra-modal dataset are contained in a multi-modal training dataset. The one or more computer processors estimate an inter-modal label utilizing inter-modal label transformation of a subsequent sample. The one or more computer processors estimate an intra-modal label utilizing intra-modal label transformation of the subsequent sample. The one or more computer processors label the subsequent sample with a cross-modal label by combining the estimated inter-modal label and the estimated intra-modal label.


