Multimedia Categorization via Multi-Modal Semantic Feature Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video categorization technologies rely solely on image modality data, disregarding text and voice modalities, leading to low accuracy in multimedia platforms, and fail to effectively utilize multi-modality interactions and semantic alignment in temporal dimensions.

Innovation Solution

A method that determines semantically relevant feature combinations by synthesizing features from multiple modalities (text, image, voice) using a computational model, incorporating an attention mechanism to improve categorization accuracy by aligning and combining features across different modalities, thereby enhancing expression capability and categorization performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video categorization uses only image modality data, then the system complexity is low, but the categorization accuracy is low

Engineering Contradiction:
Improvecategorization accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple modality information (image, text, voice) into a unified categorization framework. The computational model integrates features from different modalities through feature combination and semantic alignment, merging previously separate processing streams into a cohesive multi-modal system that achieves higher categorization accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite feature representation by combining features from different modalities (image features, text features, voice features) into a unified feature vector. This composite feature structure leverages the strengths of each modality type, analogous to how composite materials combine different substances to achieve superior properties.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If video categorization uses multiple modality information, then the categorization accuracy is improved, but the computational complexity increases

Engineering Contradiction:
Improvecategorization accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary feature extraction and combination operations before the main categorization task. By pre-processing the multi-modal data to create combined feature representations and identifying semantically relevant features in advance, the system reduces the computational burden during actual categorization operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and focuses on semantically relevant feature combinations from the multi-modal data. Rather than processing all features equally, the system identifies and extracts the most meaningful feature combinations that contribute to categorization accuracy, discarding redundant information.

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of information

If feature combinations from different modalities are synthesized, then the expression capability is strengthened, but the difficulty of detecting and measuring semantic relevance increases

Engineering Contradiction:
Improveexpression capabilityVSAvoidsemantic relevance detection
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent replaces manual or rule-based semantic relevance detection with a computational model that automatically learns semantic relationships. The model uses machine learning algorithms to identify semantically relevant feature combinations without requiring explicit programming of semantic rules, substituting mechanical detection with intelligent computation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates feedback mechanisms where the computational model learns from the interactions between different modality features. By analyzing which feature combinations lead to accurate categorization, the model refines its understanding of semantic relevance, creating a feedback loop that improves detection capability over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11675827B2Multimedia file categorizing, information processing, and model training method, system, and device
Publication Date: 2023.06.13 ALIBABA GROUP HOLDING LTD
  • US11675827B2 patent drawing
  • US11675827B2 patent drawing
  • US11675827B2 patent drawing

AI summary

Disclosed embodiments provide a multimedia file categorizing, information processing, and model training method, system, and device. The method comprises determining a plurality of feature combinations according to respective feature sets corresponding to at least two types of modality information in a multimedia file; determining a semantically relevant feature combination by using a first computational model according to the plurality of feature combinations; and categorizing the multimedia file by using the first computational model with reference to the semantically relevant feature combination. The technical solutions provided by the disclosed embodiments identify a semantically relevant feature combination in a process of categorizing a multimedia file by synthesizing features corresponding to a plurality of modalities of the multimedia file. The semantically relevant feature combination has a stronger expression capability and higher value. Categorizing the multimedia file by using this feature combination can effectively improve the categorization accuracy of the multimedia file.