Multimodal Theme Classification via Knowledge Base Entity Linking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for video theme classification, such as manual marking and single-modality machine learning, are inefficient, prone to errors, and fail to accurately classify videos in complex scenarios, especially when dealing with large multimedia databases.

Innovation Solution

A multimodality theme classification method that combines text and non-text information, including visual and audio features, using a pre-established knowledge base for entity linking and feature extraction to achieve more accurate theme classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual marking is used for video theme classification, then classification accuracy can be maintained, but efficiency and cost are significantly reduced

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual marking (mechanical human operation) with an automated machine learning system that processes video content through multiple modalities (visual, audio, text). The system automatically extracts features from video frames, audio signals, and subtitles, then classifies themes without human intervention, thereby maintaining accuracy while dramatically improving efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables videos to classify themselves by automatically extracting and processing their own multimodal features. The video content (visual frames, audio, text) serves its own classification purpose through automated feature extraction and theme determination, eliminating the need for external manual marking while maintaining classification quality.

Inventive Principle:
Principle #25Self-service

2Productivity

If single-modality machine learning is used for video classification, then processing speed is improved, but classification accuracy deteriorates in complex scenarios

Engineering Contradiction:
Improveprocessing speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges multiple modalities (visual information from video frames, audio information from sound tracks, and text information from subtitles) into a unified classification system. By combining features from all three modalities through feature fusion and joint processing, the system achieves both high processing speed and high classification accuracy in complex scenarios, overcoming the limitations of single-modality approaches.

Inventive Principle:
Principle #5Merging (Combining)

3Device complexity

If traditional classification methods are used, then implementation simplicity is maintained, but adaptability to complex scenarios is reduced

Engineering Contradiction:
Improvesystem simplicityVSAvoidscenario adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal classification system that handles multiple video scenarios (complex and simple) through a single multimodal framework. The system can process various video types (news, entertainment, education) and scenarios by adaptively weighting and fusing features from different modalities, making it versatile across diverse applications while maintaining a unified implementation architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11995117B2Theme classification method based on multimodality, device, and storage medium
Publication Date: 2024.05.28 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11995117B2 patent drawing
  • US11995117B2 patent drawing
  • US11995117B2 patent drawing

AI summary

A theme classification method based on multimodality is related to a field of a knowledge map. The method includes obtaining text information and non-text information of an object to be classified. The non-text information includes at least one of visual information and audio information. The method also includes determining an entity set of the text information based on a pre-established knowledge base, and then extracting a text feature of the object based on the text information and the entity set. The method also includes determining a theme classification of the object based on the text feature and a non-text feature of the object.