Cross-Modal Pre-Trained Model Training via Unified Representation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing pre-trained models are limited to single-modal data processing, such as text or images, and struggle to effectively handle and integrate information from multiple modalities, which is essential for comprehensive artificial intelligence systems.
Innovation Solution
A method and apparatus for acquiring a cross-modal pre-trained model by using training data that includes both single-modal and multi-modal language materials. The model performs multi-task training, incorporating cross-modal contrastive learning and single-modal learning tasks to enhance semantic comprehension and generalizable representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing pre-training methods are used for single-modal scenarios, then the model can process single-modal data effectively, but the model cannot effectively process information in various modalities
Solution Approach 1:
The patent applies multi-functionality by designing a pre-trained model that can handle both single-modal and cross-modal tasks. The model architecture is extended to process multiple modalities (text, image, audio) simultaneously while maintaining the capability to handle individual modalities independently through a unified representation space that accommodates diverse input types
Solution Approach 2:
The patent segments the training process into distinct tasks: single-modal learning tasks for each individual modality and cross-modal contrastive learning tasks for multi-modal integration. This segmentation allows the model to learn modality-specific features separately while also learning to integrate them, resolving the contradiction between specialized single-modal performance and general multi-modal capability
2Adaptability or versatility
If a unified model processes multiple modalities, then the model can handle various modalities, but the semantic comprehension capability deteriorates
Solution Approach 1:
The patent introduces a unified representation space as an intermediary that maps different modalities (text, image, audio) into a common semantic space. This intermediary layer enables the model to process multiple modalities while preserving semantic relationships, as the representation space acts as a mediator that maintains semantic integrity during cross-modal transformation
Solution Approach 2:
The patent changes the parameter space by transforming diverse modalities into a unified representation format. By projecting different modalities into a common vector space with consistent dimensional parameters, the model can process various modalities while maintaining precise semantic comprehension through standardized parameter representations
3Adaptability or versatility
If cross-modal contrastive learning is performed, then the similarity between different modalities is maximized, but the training complexity increases
Solution Approach 1:
The patent merges multiple learning objectives into a unified training framework that combines single-modal learning losses with cross-modal contrastive learning losses. By integrating these objectives into a single optimization process with a composite loss function, the model achieves cross-modal alignment without requiring separate complex training procedures for each task
Data Source
AI summary
The present disclosure discloses a method and apparatus for acquiring a pre-trained model, and relates to natural language processing and deep learning technologies in the field of artificial intelligence technologies. An implementation includes: acquiring training data, the training data including a single-modal language material and a multi-modal language material, and the multi-modal language material including a language material pair formed by a first-modal language material and a second-modal language material; and performing a multi-task training operation on a pre-trained model using the training data, the multi-task including at least one cross-modal contrastive learning task and at least one single-modal learning task; the pre-trained language model obtained in the present disclosure may learn from different forms of language materials, i.e., the single-modal language material and the multi-modal language material, such that the pre-trained language model may effectively process information in various modals.


