Audio Signal Distortion Estimation via Noise Frame Spectral Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing methods fail to accurately determine distortion in noise-contaminated voices, especially when the noise is significantly different from the sample noise, and struggle to measure distortion in noise sections during processing like directional sound reception and noise suppression, leading to inaccurate estimation results.
Innovation Solution
An audio signal processing estimating program that sets frames for voice and noise sections, calculates spectra, adjusts levels to equalize noise frames, estimates a noise model spectrum, and calculates distortion amounts based on selected frequencies to improve estimation precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If audio signal processing such as directional sound reception processing or noise suppressing processing is executed on a noise-contaminated audio signal, then noise reduction is achieved, but distortion occurs in both noise section and voice section making accurate distortion measurement difficult
Solution Approach 1:
The patent segments the audio signal into multiple frames and identifies specific voice frames and noise frames by comparing spectral characteristics. This segmentation allows separate processing and distortion measurement for voice and noise sections, resolving the contradiction by enabling accurate distortion measurement in the voice section while the noise section undergoes power reduction processing.
Solution Approach 2:
The patent applies different processing strategies to different parts of the signal: for noise frames, it calculates distortion based on spectral comparison; for voice frames, it uses noise model subtraction and selective frequency analysis. This local quality approach allows accurate distortion measurement tailored to the characteristics of each section, resolving the measurement difficulty caused by uniform processing.
2Measurement precision
If a relational expression of subjective and objective estimation values is determined using sample voice and PESQ, then estimation precision for voices contaminated with similar noise is high, but estimation precision for voices contaminated with greatly different noise is low
Solution Approach 1:
The patent changes the estimation parameters by calculating actual distortion amounts through spectral analysis and noise model subtraction rather than relying on fixed relational expressions from sample data. This allows the system to adapt to different noise types by measuring actual signal characteristics, resolving the contradiction between precision for similar noise and adaptability to different noise types.
Solution Approach 2:
The patent creates a noise model by copying and analyzing noise frames, then uses this model to subtract noise components from voice frames. This copying approach allows the system to handle various noise types by creating appropriate noise models for each case, improving both precision and adaptability.
3Object-affected harmful factors
If power is reduced in the noise section during signal processing, then noise suppression is achieved, but it becomes difficult to measure accurate distortion amount in the noise section
Solution Approach 1:
The patent performs preliminary identification of noise frames before power reduction processing by comparing spectral characteristics between input and output signals. This preliminary action allows the system to measure distortion in the noise section before processing alters the signal, or to use the identified noise characteristics to guide subsequent processing while maintaining measurement capability.
Data Source
AI summary
A computer implemented method comprising: setting a plurality of frames on a time axis between a first waveform of an input to audio processing and a second waveform of an output from the audio processing, detecting a voice frame and a noise frame in the first and second waveform, calculating a first and second spectrum from the first and second waveform, adjusting level of the first or second spectrum of the noise frame, setting the adjusted first and second spectrum of the noise frame as a third and fourth spectrum, calculating a distortion amount of the noise frame from the third and fourth spectrum, estimating a noise model spectrum from the first or second spectrum, determining a selected frequency by comparison of voice and noise frame spectrum levels, and calculating a distortion amount of the voice frame from the first and second spectrum of the voice frame at the selected frequency.


