Multimedia Content Recognition via Audio Fingerprint Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current media identification systems are ineffective in accurately identifying multimedia content, such as video, audio, and slideshows, especially when metadata or reliable sources are unavailable, due to reliance on textual data and Boolean operators, which fail to fully represent the content, especially for dynamic and newly generated content.

Innovation Solution

An automated multimedia content recognition system that clusters queries based on fingerprints, allowing for content identification without accessing the original content, by forming query clusters, comparing fingerprints, and generating identification information from consensus textual data when matches are not found in a base set of known signatures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional search methods using keywords, tags, and Boolean operators are used to identify multimedia content, then the search process is simple and fast, but the identification accuracy is poor because search terms do not fully represent the content

Engineering Contradiction:
Improvecontent identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional text-based search mechanisms with audio fingerprinting technology. Instead of using keywords and Boolean operators to search for content, the system converts audio content into unique fingerprint signatures that can be directly compared and matched. This substitution of the search mechanism fundamentally improves identification accuracy while maintaining operational simplicity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates audio fingerprints as simplified copies or representations of the original audio content. These fingerprints capture essential characteristics of the audio without requiring access to the full content, enabling accurate identification through comparison of these condensed representations rather than analyzing complete audio files or metadata.

Inventive Principle:
Principle #26Copying

2Reliability

If metadata or descriptions are attached to content sources to enable search and identification, then content can be identified through textual data, but this data is not always available or accurate for new and dynamically generated content

Engineering Contradiction:
Improvecontent identification reliabilityVSAvoidcapability to handle new content
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary audio fingerprinting of content during ingestion or initial processing, creating identification signatures before the content is published or distributed. This preliminary action ensures that even new or dynamically generated content without existing metadata can be immediately identified and searched using its pre-generated fingerprint, eliminating the need for separate metadata attachment processes.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If the system compares each query individually against known fingerprints, then the identification process is straightforward, but the processing time and computational resources increase significantly with large numbers of queries

Engineering Contradiction:
Improvequery processing efficiencyVSAvoididentification time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent merges multiple queries that share identical or similar audio fingerprints into single batch processing operations. Instead of comparing each query individually against the database of known fingerprints, the system groups queries by their fingerprint signatures and processes them collectively, dramatically reducing redundant computations and accelerating identification throughput for large volumes of similar content queries.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If the system accesses original multimedia content or reliable sources to generate identification information, then accurate content identification can be achieved, but this approach fails when content is not accessible or sources are unavailable

Engineering Contradiction:
Improvecontent identification accuracyVSAvoidsystem operability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent extracts essential identification characteristics from audio content in the form of fingerprints, separating the identification function from the need to access or store the complete original content. This extraction allows the system to achieve accurate content identification using only the compact fingerprint data, eliminating dependency on accessing original multimedia files or external reliable sources while maintaining high identification accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP3095046B1Automated multimedia content recognition
Publication Date: 2019.05.22 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3095046B1 patent drawingFigure 1
  • EP3095046B1 patent drawingFigure 2
  • EP3095046B1 patent drawingFigure 3

AI summary

An automated content recognition system accurately and reliably generates content identification information for multimedia content without accessing the multimedia content or a reliable source of the multimedia content. The system receives content-based queries having fingerprints of multimedia content. The system compares the individual queries to one another to match queries and thereby form query clusters that correspond to the same multimedia content. The system aggregates identification information from the queries in a cluster to generate reliable content identification information from otherwise unreliable identification information.