Joint Sound Model Generation for Noisy Audio Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound decomposition techniques are limited by noise and artifacts in the sound data used to generate models, leading to inefficiencies in separating desired sound sources from undesirable ones.

Innovation Solution

Joint sound model generation techniques are employed, where multiple models of sound data from different sound scenes are generated jointly by sharing learned information to encourage similar spectral components, allowing for improved separation of sound sources even in noisy environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional techniques are used to generate sound models independently from sound data, then the model generation process is simple, but the models are limited by noise and artifacts in the sound data

Engineering Contradiction:
Improvemodel qualityVSAvoidmodel generation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple independent model generation processes into a joint generation framework where models from different sound scenes are generated together. This merging allows shared information and learned representations to be utilized across multiple models, improving reliability by reducing the impact of noise and artifacts in individual sound data sets while maintaining a unified generation process that manages complexity through shared computational structures

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If sound data from noisy environments is used to generate models, then more real-world scenarios are covered, but the model accuracy deteriorates due to noise and artifacts

Engineering Contradiction:
Improvesound scene coverageVSAvoidmodel accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent converts the harmful effect of noise and artifacts in sound data into a benefit by using joint model generation across multiple sound scenes. The noise and artifacts present in individual recordings become part of a larger dataset that, when processed jointly, allows the system to learn robust representations that are invariant to specific noise patterns. This approach enables the system to handle diverse real-world scenarios while maintaining accuracy through the statistical power of aggregated data

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Manufacturing precision

If multiple models are generated independently, then each model can be optimized for its specific sound data, but information sharing between models is lost

Engineering Contradiction:
Improvemodel optimizationVSAvoidlearned information sharing
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent implements a joint model generation framework where a single generation process serves multiple sound scenes simultaneously. This universal approach allows the system to learn shared representations and information that are applicable across different sound data sets, enabling each model to benefit from the collective information while still being optimized for its specific sound scene. The multi-functional nature of the joint generation process eliminates information loss by design

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9318106B2Joint sound model generation techniques
Publication Date: 2016.04.19 ADOBE INC
  • US9318106B2 patent drawing
  • US9318106B2 patent drawing
  • US9318106B2 patent drawing

AI summary

Joint sound model generation techniques are described. In one or more implementations, a plurality of models of sound data received from a plurality of different sound scenes are jointly generated. The joint generating includes learning information as part of generating a first said model of sound data from a first one of the sound scenes and sharing the learned information for use in generating a second one of the models of sound data from a second one of the sound scenes.