Joint Sound Source Isolation and Frequency Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems separate vocal sources and determine fundamental frequencies as distinct tasks, lacking an integrated approach to jointly isolate sound sources and determine frequency data from mixed audio content.

Innovation Solution

A neural network model is employed to simultaneously isolate sound sources and determine frequency data, using configurations such as U-nets, source networks, and pitch networks, where the weights are trained to optimize both tasks jointly, enabling improved performance in karaoke and content recognition applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate systems are used for vocal source separation and fundamental frequency estimation, then each system can be optimized independently, but the overall accuracy and efficiency of audio processing deteriorates

Engineering Contradiction:
Improveaccuracy of sound source isolation and frequency determinationVSAvoidcomplexity of audio processing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges separate vocal source separation and fundamental frequency estimation systems into a single integrated neural network model. The model simultaneously processes both tasks using shared computational resources and coordinated processing, eliminating the need for separate system pipelines and improving overall accuracy through joint optimization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network model is designed with multi-functionality to perform both vocal source separation and fundamental frequency estimation simultaneously. The model uses universal processing components that handle multiple tasks, reducing system complexity while maintaining or improving performance through integrated learning.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If separate systems are used for vocal source separation and fundamental frequency estimation, then system development and optimization are simplified, but processing time and computational efficiency worsen

Engineering Contradiction:
Improveprocessing efficiency and speedVSAvoidtime required for audio processing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

By combining separate processing tasks into a single neural network model, the patent eliminates sequential processing steps and reduces total computational time. The integrated model processes both source separation and frequency estimation in parallel within a unified architecture, significantly improving processing efficiency.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network model performs preliminary processing of audio signals in a single pass, extracting both source separation and frequency information simultaneously rather than requiring multiple sequential processing stages. This preliminary action reduces overall processing time and improves throughput.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11862187B2Systems and methods for jointly estimating sound sources and frequencies from audio
Publication Date: 2024.01.02 SPOTIFY
  • US11862187B2 patent drawing
  • US11862187B2 patent drawing
  • US11862187B2 patent drawing

AI summary

An electronic device receives a first audio content item that includes a plurality of sound sources. The electronic device generates a representation of the first audio content item. The electronic device determines, from the representation of the first audio content item: a representation of an isolated sound source, and frequency data associated with the isolated sound source. Determining the representation of the isolated sound source and the frequency data associated with the isolated sound source includes using a neural network to jointly determine the representation of the isolated sound source and the frequency data associated with the isolated sound source. The electronic device determines that a portion of a second audio content item matches the first audio content item using the representation of the isolated sound source and/or the frequency data associated with the isolated sound source.