Cross-modal Processing Model for Temporal Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing visual-text models struggle with temporal semantic representations and correlations between images and videos, failing to learn temporal understanding capabilities in the pre-training stage due to limited data and high visual redundancy in video-text corpora, leading to low accuracy and efficiency in cross-modal data processing.

Innovation Solution

A cross-modal data processing method that pre-trains a processing model using concatenated image and text samples, maintaining temporal sequence correspondence and providing rich scene transition information, enabling explicit scene-level time alignment and improved learning of static and temporal information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If video-text corpora are used for pre-training, then cross-modal data processing capability is improved, but visual redundancy increases and learning efficiency decreases

Engineering Contradiction:
Improvecross-modal data processing capabilityVSAvoidvisual redundancy
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent segments video data into discrete frame sequences, where each frame is processed independently through the cross-modal processing model. This segmentation approach reduces visual redundancy by focusing on individual frame-text correlations rather than processing entire video sequences as monolithic units, thereby improving learning efficiency while maintaining cross-modal processing capability.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If existing visual-text models are used, then processing speed is maintained, but temporal semantic representation accuracy decreases

Engineering Contradiction:
Improvetemporal semantic representation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent performs preliminary feature extraction on video frames before feeding them into the cross-modal processing model. By pre-processing and organizing frame features in advance, the system achieves accurate temporal semantic representation through structured frame sequences while maintaining processing speed through efficient feature representation.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If concatenated training samples with scene transition information are used, then scene-level time alignment accuracy is improved, but training data complexity increases

Engineering Contradiction:
Improvescene-level time alignment accuracyVSAvoidtraining data complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces scene transition information as an intermediary element that bridges video frames and text descriptions. This intermediary provides explicit temporal cues that improve scene-level time alignment accuracy while managing training data complexity through structured annotation formats that capture temporal relationships efficiently.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240420458A1Cross-modal data processing method and apparatus, device, medium, and program product
Publication Date: 2024.12.19 LEMON INC(GB)
  • US20240420458A1 patent drawing
  • US20240420458A1 patent drawing
  • US20240420458A1 patent drawing

AI summary

The disclosure provides a cross-modal data processing method and apparatus, a device, a storage medium, and a program product. The method comprises: obtaining first modal data to be processed; obtaining a first modal data feature by performing feature extraction based on the first modal data; and obtaining second modal data based on the first modal data feature and a cross-modal processing model, the first modal data and the second modal data having different modalities, wherein the cross-modal processing model needs to be pre-trained based on a concatenated training sample, and the concatenated training sample comprises a concatenated image sample and a corresponding concatenated text sample.