Audio Classifier Trained via Video Frame Image Labels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio classifiers require manual annotation of audio data, making them cost-ineffective and time-consuming, and they struggle to efficiently utilize unannotated audio data from video content.

Innovation Solution

An audio classifier is trained using video frames with associated image labels, leveraging the correlation between audio and video modalities, allowing for the use of unannotated audio data and sophisticated image detection systems to quickly train the classifier.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional audio classifiers use manual annotation of audio data, then classification accuracy can be achieved, but the process becomes cost-ineffective and time-consuming

Engineering Contradiction:
Improveclassification accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses video frames as an intermediary to transfer labeling information from the visual domain to the audio domain. Image recognition models automatically generate labels for video frames, which then serve as training labels for audio segments without requiring manual audio annotation. This intermediary approach enables automated training while maintaining accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the manual mechanical process of audio annotation with an automated computational system. Instead of humans manually labeling audio data, the system uses image recognition models to generate labels automatically, substituting human labor with computational processes that are faster and more scalable.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If manual annotation is used for training audio classifiers, then labeled training data can be obtained, but the process becomes cost-ineffective

Engineering Contradiction:
Improvelabeled training dataVSAvoidtraining cost
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

Video frames serve as an intermediary that bridges the gap between unannotated audio data and labeled training data. The image recognition model processes video frames to generate labels, which are then paired with corresponding audio segments, creating labeled training data without direct manual audio annotation and reducing costs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service labeling where the training data is automatically generated through the correlation between video frames and audio segments. The image recognition model and video-audio pairing process automatically create labeled datasets without requiring human annotators, making the system self-sufficient and cost-effective.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If conventional audio classifiers are trained with manually annotated data, then they can classify audio accurately, but they cannot efficiently utilize unannotated audio data from video content

Engineering Contradiction:
Improveaudio classification accuracyVSAvoidutilization of unannotated data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal training framework that can handle both annotated and unannotated audio data. By using video frames as an intermediary, the system can process unannotated audio-video pairs and generate labels automatically, making the training process versatile and adaptable to different data types without requiring separate processes for annotated and unannotated data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Video frames act as a universal intermediary that enables the system to process unannotated audio data. The image recognition model generates labels for video frames, which then serve as training signals for corresponding audio segments, allowing the system to efficiently utilize previously unusable unannotated audio data from video content.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10566009B1Audio classifier
Publication Date: 2020.02.18 GOOGLE LLC
  • US10566009B1 patent drawing
  • US10566009B1 patent drawing
  • US10566009B1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for audio classifiers. In one aspect, a method includes obtaining a plurality of video frames from a plurality of videos, wherein each of the plurality of video frames is associated with one or more image labels of a plurality of image labels determined based on image recognition; obtaining a plurality of audio segments corresponding to the plurality of video frames, wherein each audio segment has a specified duration relative to the corresponding video frame; and generating an audio classifier trained using the plurality of audio segment and the associated image labels as input, wherein the audio classifier is trained such that the one or more groups of audio segments are determined to be associated with respective one or more audio labels.