Cross-Modal Pre-Trained Model Training via Unified Representation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing pre-trained models are limited to single-modal data processing, such as text or images, and struggle to effectively handle and integrate information from multiple modalities, which is essential for comprehensive artificial intelligence systems.

Innovation Solution

A method and apparatus for acquiring a cross-modal pre-trained model by using training data that includes both single-modal and multi-modal language materials. The model performs multi-task training, incorporating cross-modal contrastive learning and single-modal learning tasks to enhance semantic comprehension and generalizable representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing pre-training methods are used for single-modal scenarios, then the model can process single-modal data effectively, but the model cannot effectively process information in various modalities

Engineering Contradiction:
Improvemulti-modal processing capabilityVSAvoidsingle-modal processing performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies multi-functionality by designing a pre-trained model that can handle both single-modal and cross-modal tasks. The model architecture is extended to process multiple modalities (text, image, audio) simultaneously while maintaining the capability to handle individual modalities independently through a unified representation space that accommodates diverse input types

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the training process into distinct tasks: single-modal learning tasks for each individual modality and cross-modal contrastive learning tasks for multi-modal integration. This segmentation allows the model to learn modality-specific features separately while also learning to integrate them, resolving the contradiction between specialized single-modal performance and general multi-modal capability

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If a unified model processes multiple modalities, then the model can handle various modalities, but the semantic comprehension capability deteriorates

Engineering Contradiction:
Improvemulti-modal information processingVSAvoidsemantic comprehension capability
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a unified representation space as an intermediary that maps different modalities (text, image, audio) into a common semantic space. This intermediary layer enables the model to process multiple modalities while preserving semantic relationships, as the representation space acts as a mediator that maintains semantic integrity during cross-modal transformation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter space by transforming diverse modalities into a unified representation format. By projecting different modalities into a common vector space with consistent dimensional parameters, the model can process various modalities while maintaining precise semantic comprehension through standardized parameter representations

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If cross-modal contrastive learning is performed, then the similarity between different modalities is maximized, but the training complexity increases

Engineering Contradiction:
Improvecross-modal similarity alignmentVSAvoidtraining operation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple learning objectives into a unified training framework that combines single-modal learning losses with cross-modal contrastive learning losses. By integrating these objectives into a single optimization process with a composite loss function, the model achieves cross-modal alignment without requiring separate complex training procedures for each task

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12277401B2Method and apparatus for acquiring pre-trained model
Publication Date: 2025.04.15 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12277401B2 patent drawing
  • US12277401B2 patent drawing
  • US12277401B2 patent drawing

AI summary

The present disclosure discloses a method and apparatus for acquiring a pre-trained model, and relates to natural language processing and deep learning technologies in the field of artificial intelligence technologies. An implementation includes: acquiring training data, the training data including a single-modal language material and a multi-modal language material, and the multi-modal language material including a language material pair formed by a first-modal language material and a second-modal language material; and performing a multi-task training operation on a pre-trained model using the training data, the multi-task including at least one cross-modal contrastive learning task and at least one single-modal learning task; the pre-trained language model obtained in the present disclosure may learn from different forms of language materials, i.e., the single-modal language material and the multi-modal language material, such that the pre-trained language model may effectively process information in various modals.