Cross-Modality Semantic Model Training for Text Image Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multimodal processing methods fail to capture sufficient semantic information during model training, particularly in the relationship between text and vision modalities, resulting in poor training and recognition effects.

Innovation Solution

A cross-modality processing method that combines corpus and image data to generate training samples, trains a semantic model to learn semantic vectors containing combinations of both, and applies this model for cross-modality processing, enabling improved semantic relation learning and recognition between text and images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If current multimodal processing methods are used, then processing speed is maintained, but semantic information capture is insufficient and training effect deteriorates

Engineering Contradiction:
Improvesemantic informationVSAvoidtraining effect
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent merges text and image modalities into a unified multimodal processing framework. The system combines text features and image features through feature fusion layers, creating integrated representations that capture semantic relationships between different modalities. This merging enables the model to simultaneously process and understand both text and image information, resolving the contradiction by ensuring comprehensive semantic information capture while maintaining reliable training effects through unified loss functions and joint optimization

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If semantic relations between text and vision modalities are not established, then model complexity is reduced, but recognition capability deteriorates

Engineering Contradiction:
Improverecognition capabilityVSAvoidmodel structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces semantic relation modules as intermediary components that bridge text and vision modalities. These modules include attention mechanisms and relation extraction layers that explicitly model the semantic connections between text elements and image regions. The intermediaries transform complex cross-modal relationships into structured representations, improving recognition capability while managing model complexity through modular architecture and selective feature interaction

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If comprehensive multimodal training is implemented, then semantic understanding is improved, but training time increases

Engineering Contradiction:
Improvesemantic informationVSAvoidtraining time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements preliminary action through pre-training strategies and feature extraction optimizations. The system pre-extracts features from large datasets using efficient encoders, pre-computes attention matrices, and uses progressive training approaches where simpler modalities are trained first before introducing complex cross-modal interactions. This preliminary preparation reduces the computational burden during full multimodal training, enabling comprehensive semantic understanding while controlling training time through efficient data loading, caching, and parallel processing techniques

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11341366B2Cross-modality processing method and apparatus, and computer storage medium
Publication Date: 2022.05.24 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11341366B2 patent drawing
  • US11341366B2 patent drawing
  • US11341366B2 patent drawing

AI summary

A cross-modality processing method is related to a field of natural language processing technologies. The method includes: obtaining a sample set, wherein the sample set includes a plurality of corpus and a plurality of images; generating a plurality of training samples according to the sample set, in which each of the plurality of the training samples is a combination of at least one of the plurality of the corpus and at least one of the plurality of the images corresponding to the at least one of the plurality of the corpus; adopting the plurality of the training samples to train a semantic model, so that the semantic model learns semantic vectors containing combinations of the corpus and the images.