Content Processing Apparatus Using Multi-Modal Fingerprint Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing content recognition techniques fail to accurately distinguish between different contents that share the same scene, as they rely solely on visual or audio characteristic information, leading to misidentification when multiple sources provide similar content.
Innovation Solution
A content processing apparatus and method that extracts both video and audio characteristic information from content, transmitting these to a server for matching, allowing for precise recognition by filtering through a database of pre-stored information, thereby enhancing accuracy in identifying the specific content being displayed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If only visual characteristic information is used for content recognition, then the recognition process is simple and fast, but different contents with the same scene cannot be distinguished
Solution Approach 1:
The patent combines visual characteristic information and audio characteristic information into a unified content recognition system. The audio characteristic information is extracted from the audio track and matched with visual scene data, creating a multi-dimensional fingerprint that uniquely identifies content even when visual scenes are identical across different sources.
Solution Approach 2:
The patent adds an audio dimension to the traditional visual-only content recognition. By extracting audio features (such as dialogue, sound effects, music) and combining them with visual features, the system creates a multi-dimensional characteristic space that enables differentiation of contents that would be indistinguishable using only visual information.
2Measurement precision
If multiple types of characteristic information are extracted and processed, then content recognition accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary extraction of both visual and audio characteristic information and stores them as pre-processed fingerprints. This advance preparation allows for rapid matching during actual content recognition operations, reducing real-time processing requirements while maintaining high accuracy through comprehensive feature analysis.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
A content processing apparatus is provided. The content processing apparatus includes output circuitry configured to output a content, communication circuitry configured to communicate with a server, and a processor configured to extract, from the content, first characteristic information and second characteristic information of a different type from the first characteristic information, to control the communication circuitry to transmit the extracted first characteristic information to the server, and in response to receiving a plurality of matching information corresponding to the transmitted first characteristic information, to control the output circuitry to output matching information corresponding to the second characteristic information extracted from the content.