PSD Optimization for Sound Source Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source enhancement techniques, such as those described in NPL 1, often result in insufficient quality of target sound due to imprecise power spectral density (PSD) estimation, leading to deleted auditorily important components, remaining interference and background noise, and nonlinear distortion, particularly in varying noise environments.
Innovation Solution
A PSD optimization device is introduced, which includes a PSD updating unit that takes input values for target sound, interference noise, and background noise PSDs and solves an optimization problem defined by a cost function with constraints related to the frequency, temporal, and spatial structures of the sound source to generate optimized PSD output values, improving sound source enhancement capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If simple low-computation-amount PSD estimation is used for real-time processing, then processing speed is improved, but PSD estimation precision deteriorates
Solution Approach 1:
The PSD estimation problem is segmented into three separate estimation tasks for target sound, interference noise, and background noise. Each component is estimated independently using dedicated enhancement signals from beamformers oriented toward different directions, allowing precise estimation of each component while maintaining real-time processing capability through modular computation.
Solution Approach 2:
Enhancement signals are pre-computed using beamforming with different orientation directions before PSD estimation. The beamformer computes enhanced signals for target sound direction, interference noise directions, and background noise directions in advance, which are then used as inputs for respective PSD estimations, improving overall estimation precision without increasing real-time processing burden.
2Adaptability or versatility
If general-use sound source enhancement is designed without limiting settings, then adaptability is improved, but sound source enhancement capabilities deteriorate in specific applications
Solution Approach 1:
The system dynamically adapts to different usage scenarios by allowing configuration of the number of interference noise components and their directional characteristics. The beamformer can be configured with different orientation directions based on the specific application environment, enabling the same basic architecture to optimize performance for various sound source enhancement scenarios while maintaining adaptability.
Data Source
AI summary
Sound source enhancement technology is provided that is capable of improving sound source enhancement capabilities in accordance with settings of usage and applications. A PSD optimization device includes a PSD updating unit that takes a target sound PSD input value {circumflex over ( )}φS(ω, τ), an interference noise PSD input value {circumflex over ( )}φIN(ω, τ), and a background noise PSD input value {circumflex over ( )}φBN(ω, τ) as input, and generates a target sound PSD output value φS(ω, τ), an interference noise PSD output value {circumflex over ( )}φIN(ω, τ), and a background noise PSD output value {circumflex over ( )}φBN(ω, τ), by solving an optimization problem for a cost function relating to a variable uS representing a target sound PSD, a variable uIN representing an interference noise PSD, and a variable uBN representing a background noise PSD. The optimization problem for the cost function is defined using at least one of a constraint relating to a frequency structure of a sound source or a convex cost term relating to the frequency structure of the sound source, a constraint relating to a temporal structure of the sound source or a convex cost term relating to the temporal structure of the sound source, and a constraint relating to a spatial structure of the sound source or a convex cost term relating to the spatial structure of the sound source.


