Video Adapter for Temporal Feature Extraction in Multimodal Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video language pre-training methods face challenges in achieving high accuracy and efficiency for video processing tasks due to the limited availability of high-quality video data and differences between image and video data processing, leading to poor performance when transferring image-language pre-training models to video processing.

Innovation Solution

A video processing method that extracts temporal image features with temporal information from video data, enhancing the characterization capability of image features through a video adapter incorporating temporal networks and dynamic convolutional networks, and performing multi-modal fusion using cross-attention mechanisms to improve alignment and representation of video and language features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If image-language pre-training methods are used for video processing, then training data availability is improved, but processing accuracy deteriorates due to differences between image and video data

Engineering Contradiction:
Improvetraining data availabilityVSAvoidprocessing accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments video data into sequential image frames while preserving temporal relationships. The video processing model divides video input into multiple frames, extracts features from each frame, and then fuses these features with temporal information to maintain both data availability and processing accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a video adapter as an intermediary component that bridges image-language pre-training and video processing. This adapter incorporates temporal networks and dynamic convolutional networks to process temporal information from video frames, enabling the transfer of image-language knowledge to video tasks while accounting for temporal dynamics.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If temporal information is extracted and integrated into image features, then characterization capability is improved, but model complexity increases

Engineering Contradiction:
Improvecharacterization capabilityVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges temporal feature extraction with the existing image-language pre-training architecture. The temporal network and dynamic convolutional network are integrated into the video adapter, which combines temporal information from multiple frames with spatial features from individual frames, achieving enhanced characterization without requiring a completely separate complex system.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary extraction of temporal information from video frames before fusion with image features. The temporal network pre-processes temporal relationships by extracting temporal features from sequential frames, and the dynamic convolutional network prepares spatial-temporal features for subsequent fusion, organizing complex temporal data in advance to simplify later processing stages.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If video-specific pre-training is conducted to improve accuracy, then processing accuracy is improved, but training costs increase due to limited high-quality video data

Engineering Contradiction:
Improveprocessing accuracyVSAvoidtraining costs
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent makes the pre-training model universal by enabling it to handle both image and video tasks. The video adapter allows the same base model to process images and videos by adding temporal processing capabilities, eliminating the need for separate pre-training systems and reducing overall training costs while maintaining high accuracy for video-specific tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes key parameters in the pre-training process by introducing temporal dimensions. Instead of training on static images alone, the model processes sequences of frames with temporal relationships, adjusting learning parameters and feature extraction mechanisms to accommodate video-specific characteristics while leveraging existing image-language pre-training foundations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240395061A1Video processing method, apparatus, device, medium, and program product
Publication Date: 2024.11.28 LEMON INC(GB)
  • US20240395061A1 patent drawing
  • US20240395061A1 patent drawing
  • US20240395061A1 patent drawing

AI summary

The present disclosure provides a video processing method, apparatus, device, storage medium, and program product. The method includes: acquiring video data; obtaining, based on the video data, a temporal image feature with temporal information; determining, based on the temporal image feature, a target text feature in a set of text features that matches the temporal image feature; and obtaining, based on the target text feature, target text data corresponding to the video data.