Media Classification Using ML for Music Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current media content identification processes are time-consuming and costly due to the need to review or analyze all user-generated content against registered works, leading to increased computing resource costs and delayed processing times, especially on large media sharing platforms with billions of transactions.

Innovation Solution

Implementing a tiered transaction request processing scheme using machine learning models to classify media content items, where fewer resources are used for certain classifications, allowing for early termination of non-music content analysis and focusing processing on content with music, thereby reducing bandwidth and processing demands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all media content items are evaluated to identify unauthorized copies, then identification accuracy is improved, but processing time and computing resource costs increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the media content evaluation process into two distinct stages: (1) a classification stage using machine learning models to categorize content into groups such as music, non-music, live action, animation, etc., and (2) an identification stage that performs detailed unauthorized copy detection only on specific segments (music content) that require it. This segmentation allows the system to maintain high identification accuracy for copyrighted content while avoiding unnecessary processing of content types that are less likely to contain unauthorized copies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary classification action to all media content items before performing detailed identification analysis. By first passing content through machine learning classification models that predict categories and copyright likelihood, the system prepares the content in advance by identifying which items warrant further scrutiny. This preliminary action filters out content that can be quickly dismissed, allowing resources to be focused on content that requires thorough identification analysis, thereby reducing overall processing time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If all media content items are evaluated to identify unauthorized copies, then identification accuracy is improved, but computing resource costs increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidcomputing resource costs
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the media content evaluation process into two distinct stages: (1) a classification stage using machine learning models to categorize content into groups such as music, non-music, live action, animation, etc., and (2) an identification stage that performs detailed unauthorized copy detection only on specific segments (music content) that require it. This segmentation allows the system to maintain high identification accuracy for copyrighted content while avoiding unnecessary processing of content types that are less likely to contain unauthorized copies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing complete identification analysis only on music content and partial or no analysis on other content types. The machine learning classification models provide sufficient information for many content types without requiring full identification processing. This partial action approach maintains accuracy for content where it matters most (music) while reducing computing resource costs by avoiding excessive processing of content types that can be handled with lighter-weight classification alone.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If tiered processing is implemented to reduce resource usage, then processing efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the media content evaluation process into two distinct stages: (1) a classification stage using machine learning models to categorize content into groups such as music, non-music, live action, animation, etc., and (2) an identification stage that performs detailed unauthorized copy detection only on specific segments (music content) that require it. This segmentation allows the system to maintain high identification accuracy for copyrighted content while avoiding unnecessary processing of content types that are less likely to contain unauthorized copies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary classification action to all media content items before performing detailed identification analysis. By first passing content through machine learning classification models that predict categories and copyright likelihood, the system prepares the content in advance by identifying which items warrant further scrutiny. This preliminary action filters out content that can be quickly dismissed, allowing resources to be focused on content that requires thorough identification analysis, thereby reducing overall processing time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230244710A1Media classification and identification using machine learning
Publication Date: 2023.08.03 AUDIBLE MAGIC CORP
  • US20230244710A1 patent drawing
  • US20230244710A1 patent drawing
  • US20230244710A1 patent drawing

AI summary

A method and system classify media content items and then identify a subset of the classified media content items. In embodiments, audio features of a plurality of media content items are processed by one or more machine learning model to classify each of the media content items as containing music or not containing music. Those media content items not containing music are filtered out, digital fingerprints of those media content items that contain music are generated, and the digital fingerprints are used to identify those media content items that contain music.